[
  {
    "url": "https://docs.crawl4ai.com",
    "title": "Home - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-open-source-llm-friendly-web-crawler-scraper\">🚀🤖 Crawl4AI: Open-Source LLM-Friendly Web Crawler &amp; Scraper</h1>\n<div class=\"badges\" align=\"center\">\n\n  <p>\n    <a href=\"https://trendshift.io/repositories/11716\" target=\"_blank\">\n      <img src=\"https://trendshift.io/api/badge/repositories/11716\" alt=\"unclecode%2Fcrawl4ai | Trendshift\" style=\"width: 250px; height: 55px;\" width=\"250\" height=\"55\">\n    </a>\n\n  </p>\n\n  <p>\n    <a href=\"https://github.com/unclecode/crawl4ai/stargazers\">\n      <img src=\"https://img.shields.io/github/stars/unclecode/crawl4ai?style=social\" alt=\"GitHub Stars\">\n    </a>\n    <a href=\"https://github.com/unclecode/crawl4ai/network/members\">\n      <img src=\"https://img.shields.io/github/forks/unclecode/crawl4ai?style=social\" alt=\"GitHub Forks\">\n    </a>\n    <a href=\"https://badge.fury.io/py/crawl4ai\">\n      <img src=\"https://badge.fury.io/py/crawl4ai.svg\" alt=\"PyPI version\">\n    </a>\n  </p>\n\n  <p>\n    <a href=\"https://pypi.org/project/crawl4ai/\">\n      <img src=\"https://img.shields.io/pypi/pyversions/crawl4ai\" alt=\"Python Version\">\n    </a>\n    <a href=\"https://pepy.tech/project/crawl4ai\">\n      <img src=\"https://static.pepy.tech/badge/crawl4ai/month\" alt=\"Downloads\">\n    </a>\n    <a href=\"https://github.com/unclecode/crawl4ai/blob/main/LICENSE\">\n      <img src=\"https://img.shields.io/github/license/unclecode/crawl4ai\" alt=\"License\">\n    </a>\n  </p>\n  <p align=\"center\">\n    <a href=\"https://x.com/crawl4ai\">\n      <img src=\"https://img.shields.io/badge/Follow%20on%20X-000000?style=for-the-badge&amp;logo=x&amp;logoColor=white\" alt=\"Follow on X\">\n    </a>\n    <a href=\"https://www.linkedin.com/company/crawl4ai\">\n      <img src=\"https://img.shields.io/badge/Follow%20on%20LinkedIn-0077B5?style=for-the-badge&amp;logo=linkedin&amp;logoColor=white\" alt=\"Follow on LinkedIn\">\n    </a>\n    <a href=\"https://discord.gg/jP8KfhDhyN\">\n      <img src=\"https://img.shields.io/badge/Join%20our%20Discord-5865F2?style=for-the-badge&amp;logo=discord&amp;logoColor=white\" alt=\"Join our Discord\">\n    </a>\n  </p>\n\n</div>\n\n<hr>\n<h4 id=\"crawl4ai-cloud-api-closed-beta-launching-soon\">🚀 Crawl4AI Cloud API — Closed Beta (Launching Soon)</h4>\n<p>Reliable, large-scale web extraction, now built to be <em><strong>drastically more cost-effective</strong></em> than any of the existing solutions.</p>\n<p>👉 <strong>Apply <a href=\"https://forms.gle/E9MyPaNXACnAMaqG7\">here</a> for early access</strong><br>\n<em>We’ll be onboarding in phases and working closely with early users.\nLimited slots.</em></p>\n<hr>\n<p>Crawl4AI is the #1 trending GitHub repository, actively maintained by a vibrant community. It delivers blazing-fast, AI-ready web crawling tailored for large language models, AI agents, and data pipelines. Fully open source, flexible, and built for real-time performance, <strong>Crawl4AI</strong> empowers developers with unmatched speed, precision, and deployment ease.</p>\n<blockquote>\n<p>Enjoy using Crawl4AI? Consider <strong><a href=\"https://github.com/sponsors/unclecode\">becoming a sponsor</a></strong> to support ongoing development and community growth!</p>\n</blockquote>\n<h2 id=\"ai-assistant-skill-now-available\">🆕 AI Assistant Skill Now Available!</h2>\n<div style=\"background: linear-gradient(135deg, #667eea 0%, #764ba2 100%); padding: 20px; border-radius: 10px; margin: 20px 0; box-shadow: 0 4px 6px rgba(0,0,0,0.1);\">\n  <h3 style=\"color: white; margin: 0 0 10px 0;\" id=\"toc-heading-2--crawl4ai-skill-for-claude--ai-assistants\">🤖 Crawl4AI Skill for Claude &amp; AI Assistants</h3>\n  <p style=\"color: white; margin: 10px 0;\">Supercharge your AI coding assistant with complete Crawl4AI knowledge! Download our comprehensive skill package that includes:</p>\n  <ul style=\"color: white; margin: 10px 0;\">\n    <li>📚 Complete SDK reference (23K+ words)</li>\n    <li>🚀 Ready-to-use extraction scripts</li>\n    <li>⚡ Schema generation for efficient scraping</li>\n    <li>🔧 Version 0.7.4 compatible</li>\n  </ul>\n  <div style=\"text-align: center; margin-top: 15px;\">\n    <a href=\"assets/crawl4ai-skill.zip\" download=\"\" style=\"background: white; color: #667eea; padding: 12px 30px; border-radius: 5px; text-decoration: none; font-weight: bold; display: inline-block; transition: transform 0.2s;\">\n      📦 Download Skill Package\n    </a>\n  </div>\n  <p style=\"color: white; margin: 15px 0 0 0; font-size: 0.9em; text-align: center;\">\n    Works with Claude, Cursor, Windsurf, and other AI coding assistants. Import the .zip file into your AI assistant's skill/knowledge system.\n  </p>\n</div>\n\n<h2 id=\"new-adaptive-web-crawling\">🎯 New: Adaptive Web Crawling</h2>\n<p>Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines when sufficient information has been gathered to answer your query.</p>\n<p><a href=\"core/adaptive-crawling/\">Learn more about Adaptive Crawling →</a></p>\n<h2 id=\"quick-start\">Quick Start</h2>\n<p>Here's a quick example to show you how easy it is to use Crawl4AI with its asynchronous capabilities:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Create an instance of AsyncWebCrawler</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Run the crawler on a URL</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://crawl4ai.com\"</span>)\n\n        <span class=\"hljs-comment\"># Print the extracted content</span>\n        <span class=\"hljs-built_in\">print</span>(result.markdown)\n\n<span class=\"hljs-comment\"># Run the async main function</span>\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"video-tutorial\">Video Tutorial</h2>\n<div align=\"center\">\n  \n</div>\n\n<hr>\n<h2 id=\"what-does-crawl4ai-do\">What Does Crawl4AI Do?</h2>\n<p>Crawl4AI is a feature-rich crawler and scraper that aims to:</p>\n<p>1. <strong>Generate Clean Markdown</strong>: Perfect for RAG pipelines or direct ingestion into LLMs.<br>\n2. <strong>Structured Extraction</strong>: Parse repeated patterns with CSS, XPath, or LLM-based extraction.<br>\n3. <strong>Advanced Browser Control</strong>: Hooks, proxies, stealth modes, session re-use—fine-grained control.<br>\n4. <strong>High Performance</strong>: Parallel crawling, chunk-based extraction, real-time use cases.<br>\n5. <strong>Open Source</strong>: No forced API keys, no paywalls—everyone can access their data.  </p>\n<p><strong>Core Philosophies</strong>:\n- <strong>Democratize Data</strong>: Free to use, transparent, and highly configurable.<br>\n- <strong>LLM Friendly</strong>: Minimally processed, well-structured text, images, and metadata, so AI models can easily consume it.</p>\n<hr>\n<h2 id=\"documentation-structure\">Documentation Structure</h2>\n<p>To help you get started, we’ve organized our docs into clear sections:</p>\n<ul>\n<li><strong>Setup &amp; Installation</strong><br>\n  Basic instructions to install Crawl4AI via pip or Docker.  </li>\n<li><strong>Quick Start</strong><br>\n  A hands-on introduction showing how to do your first crawl, generate Markdown, and do a simple extraction.  </li>\n<li><strong>Core</strong><br>\n  Deeper guides on single-page crawling, advanced browser/crawler parameters, content filtering, and caching.  </li>\n<li><strong>Advanced</strong><br>\n  Explore link &amp; media handling, lazy loading, hooking &amp; authentication, proxies, session management, and more.  </li>\n<li><strong>Extraction</strong><br>\n  Detailed references for no-LLM (CSS, XPath) vs. LLM-based strategies, chunking, and clustering approaches.  </li>\n<li><strong>API Reference</strong><br>\n  Find the technical specifics of each class and method, including <code>AsyncWebCrawler</code>, <code>arun()</code>, and <code>CrawlResult</code>.</li>\n</ul>\n<p>Throughout these sections, you’ll find code samples you can <strong>copy-paste</strong> into your environment. If something is missing or unclear, raise an issue or PR.</p>\n<hr>\n<h2 id=\"how-you-can-support\">How You Can Support</h2>\n<ul>\n<li><strong>Star &amp; Fork</strong>: If you find Crawl4AI helpful, star the repo on GitHub or fork it to add your own features.  </li>\n<li><strong>File Issues</strong>: Encounter a bug or missing feature? Let us know by filing an issue, so we can improve.  </li>\n<li><strong>Pull Requests</strong>: Whether it’s a small fix, a big feature, or better docs—contributions are always welcome.  </li>\n<li><strong>Join Discord</strong>: Come chat about web scraping, crawling tips, or AI workflows with the community.  </li>\n<li><strong>Spread the Word</strong>: Mention Crawl4AI in your blog posts, talks, or on social media.  </li>\n</ul>\n<p><strong>Our mission</strong>: to empower everyone—students, researchers, entrepreneurs, data scientists—to access, parse, and shape the world’s data with speed, cost-efficiency, and creative freedom.</p>\n<hr>\n<h2 id=\"quick-links\">Quick Links</h2>\n<ul>\n<li><strong><a href=\"https://github.com/unclecode/crawl4ai\">GitHub Repo</a></strong>  </li>\n<li><strong><a href=\"core/installation/\">Installation Guide</a></strong>  </li>\n<li><strong><a href=\"core/quickstart/\">Quick Start</a></strong>  </li>\n<li><strong><a href=\"api/async-webcrawler/\">API Reference</a></strong>  </li>\n<li><strong><a href=\"https://github.com/unclecode/crawl4ai/blob/main/CHANGELOG.md\">Changelog</a></strong>  </li>\n</ul>\n<p>Thank you for joining me on this journey. Let’s keep building an <strong>open, democratic</strong> approach to data extraction and AI together.</p>\n<p>Happy Crawling!<br>\n— <em>Unclecode, Founder &amp; Maintainer of Crawl4AI</em>  </p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/",
    "title": "Home - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-open-source-llm-friendly-web-crawler-scraper\">🚀🤖 Crawl4AI: Open-Source LLM-Friendly Web Crawler &amp; Scraper</h1>\n<div class=\"badges\" align=\"center\">\n\n  <p>\n    <a href=\"https://trendshift.io/repositories/11716\" target=\"_blank\">\n      <img src=\"https://trendshift.io/api/badge/repositories/11716\" alt=\"unclecode%2Fcrawl4ai | Trendshift\" style=\"width: 250px; height: 55px;\" width=\"250\" height=\"55\">\n    </a>\n\n  </p>\n\n  <p>\n    <a href=\"https://github.com/unclecode/crawl4ai/stargazers\">\n      <img src=\"https://img.shields.io/github/stars/unclecode/crawl4ai?style=social\" alt=\"GitHub Stars\">\n    </a>\n    <a href=\"https://github.com/unclecode/crawl4ai/network/members\">\n      <img src=\"https://img.shields.io/github/forks/unclecode/crawl4ai?style=social\" alt=\"GitHub Forks\">\n    </a>\n    <a href=\"https://badge.fury.io/py/crawl4ai\">\n      <img src=\"https://badge.fury.io/py/crawl4ai.svg\" alt=\"PyPI version\">\n    </a>\n  </p>\n\n  <p>\n    <a href=\"https://pypi.org/project/crawl4ai/\">\n      <img src=\"https://img.shields.io/pypi/pyversions/crawl4ai\" alt=\"Python Version\">\n    </a>\n    <a href=\"https://pepy.tech/project/crawl4ai\">\n      <img src=\"https://static.pepy.tech/badge/crawl4ai/month\" alt=\"Downloads\">\n    </a>\n    <a href=\"https://github.com/unclecode/crawl4ai/blob/main/LICENSE\">\n      <img src=\"https://img.shields.io/github/license/unclecode/crawl4ai\" alt=\"License\">\n    </a>\n  </p>\n  <p align=\"center\">\n    <a href=\"https://x.com/crawl4ai\">\n      <img src=\"https://img.shields.io/badge/Follow%20on%20X-000000?style=for-the-badge&amp;logo=x&amp;logoColor=white\" alt=\"Follow on X\">\n    </a>\n    <a href=\"https://www.linkedin.com/company/crawl4ai\">\n      <img src=\"https://img.shields.io/badge/Follow%20on%20LinkedIn-0077B5?style=for-the-badge&amp;logo=linkedin&amp;logoColor=white\" alt=\"Follow on LinkedIn\">\n    </a>\n    <a href=\"https://discord.gg/jP8KfhDhyN\">\n      <img src=\"https://img.shields.io/badge/Join%20our%20Discord-5865F2?style=for-the-badge&amp;logo=discord&amp;logoColor=white\" alt=\"Join our Discord\">\n    </a>\n  </p>\n\n</div>\n\n<hr>\n<h4 id=\"crawl4ai-cloud-api-closed-beta-launching-soon\">🚀 Crawl4AI Cloud API — Closed Beta (Launching Soon)</h4>\n<p>Reliable, large-scale web extraction, now built to be <em><strong>drastically more cost-effective</strong></em> than any of the existing solutions.</p>\n<p>👉 <strong>Apply <a href=\"https://forms.gle/E9MyPaNXACnAMaqG7\">here</a> for early access</strong><br>\n<em>We’ll be onboarding in phases and working closely with early users.\nLimited slots.</em></p>\n<hr>\n<p>Crawl4AI is the #1 trending GitHub repository, actively maintained by a vibrant community. It delivers blazing-fast, AI-ready web crawling tailored for large language models, AI agents, and data pipelines. Fully open source, flexible, and built for real-time performance, <strong>Crawl4AI</strong> empowers developers with unmatched speed, precision, and deployment ease.</p>\n<blockquote>\n<p>Enjoy using Crawl4AI? Consider <strong><a href=\"https://github.com/sponsors/unclecode\">becoming a sponsor</a></strong> to support ongoing development and community growth!</p>\n</blockquote>\n<h2 id=\"ai-assistant-skill-now-available\">🆕 AI Assistant Skill Now Available!</h2>\n<div style=\"background: linear-gradient(135deg, #667eea 0%, #764ba2 100%); padding: 20px; border-radius: 10px; margin: 20px 0; box-shadow: 0 4px 6px rgba(0,0,0,0.1);\">\n  <h3 style=\"color: white; margin: 0 0 10px 0;\" id=\"toc-heading-2--crawl4ai-skill-for-claude--ai-assistants\">🤖 Crawl4AI Skill for Claude &amp; AI Assistants</h3>\n  <p style=\"color: white; margin: 10px 0;\">Supercharge your AI coding assistant with complete Crawl4AI knowledge! Download our comprehensive skill package that includes:</p>\n  <ul style=\"color: white; margin: 10px 0;\">\n    <li>📚 Complete SDK reference (23K+ words)</li>\n    <li>🚀 Ready-to-use extraction scripts</li>\n    <li>⚡ Schema generation for efficient scraping</li>\n    <li>🔧 Version 0.7.4 compatible</li>\n  </ul>\n  <div style=\"text-align: center; margin-top: 15px;\">\n    <a href=\"assets/crawl4ai-skill.zip\" download=\"\" style=\"background: white; color: #667eea; padding: 12px 30px; border-radius: 5px; text-decoration: none; font-weight: bold; display: inline-block; transition: transform 0.2s;\">\n      📦 Download Skill Package\n    </a>\n  </div>\n  <p style=\"color: white; margin: 15px 0 0 0; font-size: 0.9em; text-align: center;\">\n    Works with Claude, Cursor, Windsurf, and other AI coding assistants. Import the .zip file into your AI assistant's skill/knowledge system.\n  </p>\n</div>\n\n<h2 id=\"new-adaptive-web-crawling\">🎯 New: Adaptive Web Crawling</h2>\n<p>Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines when sufficient information has been gathered to answer your query.</p>\n<p><a href=\"core/adaptive-crawling/\">Learn more about Adaptive Crawling →</a></p>\n<h2 id=\"quick-start\">Quick Start</h2>\n<p>Here's a quick example to show you how easy it is to use Crawl4AI with its asynchronous capabilities:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Create an instance of AsyncWebCrawler</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Run the crawler on a URL</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://crawl4ai.com\"</span>)\n\n        <span class=\"hljs-comment\"># Print the extracted content</span>\n        <span class=\"hljs-built_in\">print</span>(result.markdown)\n\n<span class=\"hljs-comment\"># Run the async main function</span>\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"video-tutorial\">Video Tutorial</h2>\n<div align=\"center\">\n  \n</div>\n\n<hr>\n<h2 id=\"what-does-crawl4ai-do\">What Does Crawl4AI Do?</h2>\n<p>Crawl4AI is a feature-rich crawler and scraper that aims to:</p>\n<p>1. <strong>Generate Clean Markdown</strong>: Perfect for RAG pipelines or direct ingestion into LLMs.<br>\n2. <strong>Structured Extraction</strong>: Parse repeated patterns with CSS, XPath, or LLM-based extraction.<br>\n3. <strong>Advanced Browser Control</strong>: Hooks, proxies, stealth modes, session re-use—fine-grained control.<br>\n4. <strong>High Performance</strong>: Parallel crawling, chunk-based extraction, real-time use cases.<br>\n5. <strong>Open Source</strong>: No forced API keys, no paywalls—everyone can access their data.  </p>\n<p><strong>Core Philosophies</strong>:\n- <strong>Democratize Data</strong>: Free to use, transparent, and highly configurable.<br>\n- <strong>LLM Friendly</strong>: Minimally processed, well-structured text, images, and metadata, so AI models can easily consume it.</p>\n<hr>\n<h2 id=\"documentation-structure\">Documentation Structure</h2>\n<p>To help you get started, we’ve organized our docs into clear sections:</p>\n<ul>\n<li><strong>Setup &amp; Installation</strong><br>\n  Basic instructions to install Crawl4AI via pip or Docker.  </li>\n<li><strong>Quick Start</strong><br>\n  A hands-on introduction showing how to do your first crawl, generate Markdown, and do a simple extraction.  </li>\n<li><strong>Core</strong><br>\n  Deeper guides on single-page crawling, advanced browser/crawler parameters, content filtering, and caching.  </li>\n<li><strong>Advanced</strong><br>\n  Explore link &amp; media handling, lazy loading, hooking &amp; authentication, proxies, session management, and more.  </li>\n<li><strong>Extraction</strong><br>\n  Detailed references for no-LLM (CSS, XPath) vs. LLM-based strategies, chunking, and clustering approaches.  </li>\n<li><strong>API Reference</strong><br>\n  Find the technical specifics of each class and method, including <code>AsyncWebCrawler</code>, <code>arun()</code>, and <code>CrawlResult</code>.</li>\n</ul>\n<p>Throughout these sections, you’ll find code samples you can <strong>copy-paste</strong> into your environment. If something is missing or unclear, raise an issue or PR.</p>\n<hr>\n<h2 id=\"how-you-can-support\">How You Can Support</h2>\n<ul>\n<li><strong>Star &amp; Fork</strong>: If you find Crawl4AI helpful, star the repo on GitHub or fork it to add your own features.  </li>\n<li><strong>File Issues</strong>: Encounter a bug or missing feature? Let us know by filing an issue, so we can improve.  </li>\n<li><strong>Pull Requests</strong>: Whether it’s a small fix, a big feature, or better docs—contributions are always welcome.  </li>\n<li><strong>Join Discord</strong>: Come chat about web scraping, crawling tips, or AI workflows with the community.  </li>\n<li><strong>Spread the Word</strong>: Mention Crawl4AI in your blog posts, talks, or on social media.  </li>\n</ul>\n<p><strong>Our mission</strong>: to empower everyone—students, researchers, entrepreneurs, data scientists—to access, parse, and shape the world’s data with speed, cost-efficiency, and creative freedom.</p>\n<hr>\n<h2 id=\"quick-links\">Quick Links</h2>\n<ul>\n<li><strong><a href=\"https://github.com/unclecode/crawl4ai\">GitHub Repo</a></strong>  </li>\n<li><strong><a href=\"core/installation/\">Installation Guide</a></strong>  </li>\n<li><strong><a href=\"core/quickstart/\">Quick Start</a></strong>  </li>\n<li><strong><a href=\"api/async-webcrawler/\">API Reference</a></strong>  </li>\n<li><strong><a href=\"https://github.com/unclecode/crawl4ai/blob/main/CHANGELOG.md\">Changelog</a></strong>  </li>\n</ul>\n<p>Thank you for joining me on this journey. Let’s keep building an <strong>open, democratic</strong> approach to data extraction and AI together.</p>\n<p>Happy Crawling!<br>\n— <em>Unclecode, Founder &amp; Maintainer of Crawl4AI</em>  </p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/CONTRIBUTING/",
    "title": "Contributing Guide - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"contributing-to-crawl4ai\">Contributing to Crawl4AI</h1>\n<p>Welcome to the Crawl4AI project! As an open-source library for web crawling and AI integration, we value contributions from the community. This guide explains our branching strategy, how to contribute effectively, and the overall release process. Our goal is to maintain a stable, collaborative environment where bug fixes, features, and improvements can be integrated smoothly while allowing for experimental development.</p>\n<p>We follow a GitFlow-inspired workflow to ensure predictability and quality. Releases occur approximately every two weeks, with a focus on semantic versioning, comprehensive documentation, and user-friendly updates.</p>\n<h2 id=\"core-branches\">Core Branches</h2>\n<ul>\n<li><strong>main</strong>: The stable branch containing production-ready code. It's always identical to the latest released version and is tagged for releases. Do not submit PRs directly here.</li>\n<li><strong>develop</strong>: The primary integration branch for ongoing development. This is where all contributions (bug fixes, minor features, documentation updates) are merged. Submit your pull requests targeting this branch.</li>\n<li><strong>next</strong>: Reserved for the lead maintainer (Unclecode) to experiment with major features, refactors, or cutting-edge changes. These are merged into <code>develop</code> when ready.</li>\n<li><strong>release/vX.Y.Z</strong>: Temporary branches created from <code>develop</code> for final release preparations (e.g., version bumps, demos, release notes). These are short-lived and deleted after the release.</li>\n</ul>\n<h2 id=\"contributor-workflow\">Contributor Workflow</h2>\n<p>We encourage contributions of all kinds: bug fixes, new features, documentation improvements, tests, or even Docker enhancements. Follow these steps to contribute:</p>\n<ol>\n<li><strong>Fork the Repository</strong>: Create your own fork on GitHub.</li>\n<li>\n<p><strong>Create a Branch</strong>: Base your work on the <code>develop</code> branch.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">git checkout develop\ngit checkout -b feature/your-feature-name  <span class=\"hljs-comment\"># Or bugfix/your-bugfix-name</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n</li>\n<li>\n<p><strong>Make Changes</strong>:</p>\n<ul>\n<li>Implement your feature or fix.</li>\n<li>If updating documentation (e.g., README.md, mkdocs.yml, or docs/blog/), ensure version references are consistent (e.g., update site_name in mkdocs.yml to reflect the upcoming version if relevant).</li>\n<li>For Docker-related changes (e.g., Dockerfile, docker-compose.yml, or docs/md_v2/core/docker-deployment.md), test locally and include build instructions in your PR description.</li>\n<li>Add tests if applicable (run <code>pytest</code> to verify).</li>\n<li>Follow code style guidelines (use black for formatting).</li>\n</ul>\n</li>\n<li><strong>Commit and Push</strong>:<ul>\n<li>Use descriptive commit messages (e.g., \"Fix: Resolve issue with async crawling\").</li>\n<li>Push to your fork: <code>git push origin feature/your-feature-name</code>.</li>\n</ul>\n</li>\n<li><strong>Submit a Pull Request</strong>:<ul>\n<li>Target the <code>develop</code> branch.</li>\n<li>Provide a clear description: What does it do? Link to any issues. Include screenshots or code examples if helpful.</li>\n<li>If your change affects documentation or Docker, mention how it aligns with the version (e.g., \"Updates Docker docs for v0.7.0 compatibility\").</li>\n<li>We'll review and merge approved PRs into <code>develop</code>.</li>\n</ul>\n</li>\n<li><strong>Discuss Large Changes</strong>: For major features or experimental ideas, open an issue first to align with the project's direction.</li>\n</ol>\n<p>If your PR involves breaking changes, include a migration guide in the description.</p>\n<h2 id=\"lead-maintainers-workflow-for-reference\">Lead Maintainer's Workflow (For Reference)</h2>\n<ul>\n<li>The lead maintainer (Unclecode) uses the <code>next</code> branch for isolated experimental work.</li>\n<li>Features from <code>next</code> are periodically merged into <code>develop</code> (via rebase and merge) to keep everything in sync.</li>\n<li>This isolation ensures your contributions aren't disrupted by ongoing major changes.</li>\n</ul>\n<h2 id=\"release-process-high-level-overview\">Release Process (High-Level Overview)</h2>\n<p>Releases happen bi-weekly to ship improvements regularly. As a contributor, your merged changes in <code>develop</code> will be included in the next release unless specified otherwise. Here's a summary of what happens:</p>\n<ul>\n<li><strong>Preparation</strong>: A temporary <code>release/vX.Y.Z</code> branch is created from <code>develop</code>. Any ready features from <code>next</code> are merged here.</li>\n<li><strong>Final Updates</strong>:<ul>\n<li>Version bump in code (e.g., <code>__version__.py</code>).</li>\n<li>Creation of a demo script in <code>examples/</code> to showcase new features.</li>\n<li>Writing release notes in <code>docs/blog/</code> (personal \"I\" voice from Unclecode, with code examples, impacts, and migration guides if needed).</li>\n<li>Documentation updates: README.md (highlights, version refs), mkdocs.yml (site_name with version), docs/blog/index.md (add new release), and copying notes to <code>docs/md_v2/blog/releases/</code>.</li>\n<li>Docker updates: Dockerfile (version arg), docker-compose.yml, deploy/docker/README.md, and docs/md_v2/core/docker-deployment.md. A release candidate image (e.g., <code>X.Y.Z-r1</code>) is built and tested.</li>\n</ul>\n</li>\n<li><strong>Testing and Merge</strong>: Full tests run; changes committed and merged to <code>main</code> with a tag.</li>\n<li><strong>Publication</strong>: Tagged release on GitHub (with notes), publish to PyPI, and push Docker images (stable and <code>latest</code> after testing).</li>\n<li><strong>Sync</strong>: Back-merge to <code>develop</code> and reset <code>next</code> for the next cycle.</li>\n</ul>\n<p>Semantic versioning is used: MAJOR for breaking changes, MINOR for features, PATCH for fixes. Pre-releases (e.g., <code>-rc1</code>) may be used for testing.</p>\n<p>If your contribution requires Docker testing or affects docs, it may be part of this step—feel free to suggest updates in your PR.</p>\n<h2 id=\"benefits-of-this-approach\">Benefits of This Approach</h2>\n<ul>\n<li><strong>Stability</strong>: <code>main</code> is always reliable for users.</li>\n<li><strong>Collaboration</strong>: Fixed PR target (<code>develop</code>) makes contributing straightforward.</li>\n<li><strong>Isolation</strong>: Experimental work in <code>next</code> doesn't block team progress.</li>\n<li><strong>User-Focused</strong>: Releases include demos, detailed notes, and updated docs/Docker for easy adoption.</li>\n<li><strong>Predictability</strong>: Bi-weekly cadence keeps the project active.</li>\n</ul>\n<h2 id=\"checklist-for-contributors\">Checklist for Contributors</h2>\n<p>Before submitting a PR:</p>\n<ul>\n<li>[ ]  Based on and targeting <code>develop</code>.</li>\n<li>[ ]  Tests pass (<code>pytest</code>).</li>\n<li>[ ]  Docs updated if needed (e.g., version refs in mkdocs.yml, Docker files).</li>\n<li>[ ]  No breaking changes without a migration guide.</li>\n<li>[ ]  Descriptive title and description.</li>\n</ul>\n<h2 id=\"common-issues\">Common Issues</h2>\n<ul>\n<li><strong>Merge Conflicts</strong>: Rebase your branch on latest <code>develop</code> before PR.</li>\n<li><strong>Docker Builds</strong>: Test multi-arch (amd64/arm64) locally if changing Dockerfile.</li>\n<li><strong>Version Consistency</strong>: Ensure any version mentions match semantic rules.</li>\n</ul>\n<h2 id=\"communication\">Communication</h2>\n<ul>\n<li>Open issues for discussions or bugs.</li>\n<li>Join our Discord (link in README) for real-time help.</li>\n<li>After releases, announcements go to GitHub, Discord, and social media.</li>\n</ul>\n<p>Thanks for contributing to Crawl4AI — we appreciate your help in making it better!</p>\n<p><em>Last Updated: Feb 3, 2026</em></p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/adaptive-strategies/",
    "title": "Adaptive Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"advanced-adaptive-strategies\">Advanced Adaptive Strategies</h1>\n<h2 id=\"overview\">Overview</h2>\n<p>While the default adaptive crawling configuration works well for most use cases, understanding the underlying strategies and scoring mechanisms allows you to fine-tune the crawler for specific domains and requirements.</p>\n<h2 id=\"the-three-layer-scoring-system\">The Three-Layer Scoring System</h2>\n<h3 id=\"1-coverage-score\">1. Coverage Score</h3>\n<p>Coverage measures how comprehensively your knowledge base covers the query terms and related concepts.</p>\n<h4 id=\"mathematical-foundation\">Mathematical Foundation</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">Coverage(K, Q) = Σ(t ∈ Q) score(t, K) / |Q|\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">where</span> <span class=\"hljs-title\">score</span>(<span class=\"hljs-params\">t, K</span>)</span> = doc_coverage(t) × (<span class=\"hljs-number\">1</span> + freq_boost(t))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"components\">Components</h4>\n<ul>\n<li><strong>Document Coverage</strong>: Percentage of documents containing the term</li>\n<li><strong>Frequency Boost</strong>: Logarithmic bonus for term frequency</li>\n<li><strong>Query Decomposition</strong>: Handles multi-word queries intelligently</li>\n</ul>\n<h4 id=\"tuning-coverage\">Tuning Coverage</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># For technical documentation with specific terminology</span>\nconfig = AdaptiveConfig(\n    confidence_threshold=0.85,  <span class=\"hljs-comment\"># Require high coverage</span>\n    top_k_links=5              <span class=\"hljs-comment\"># Cast wider net</span>\n)\n\n<span class=\"hljs-comment\"># For general topics with synonyms</span>\nconfig = AdaptiveConfig(\n    confidence_threshold=0.6,   <span class=\"hljs-comment\"># Lower threshold</span>\n    top_k_links=2              <span class=\"hljs-comment\"># More focused</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-consistency-score\">2. Consistency Score</h3>\n<p>Consistency evaluates whether the information across pages is coherent and non-contradictory.</p>\n<h4 id=\"how-it-works\">How It Works</h4>\n<ol>\n<li>Extracts key statements from each document</li>\n<li>Compares statements across documents</li>\n<li>Measures agreement vs. contradiction</li>\n<li>Returns normalized score (0-1)</li>\n</ol>\n<h4 id=\"practical-impact\">Practical Impact</h4>\n<ul>\n<li><strong>High consistency (&gt;0.8)</strong>: Information is reliable and coherent</li>\n<li><strong>Medium consistency (0.5-0.8)</strong>: Some variation, but generally aligned</li>\n<li><strong>Low consistency (&lt;0.5)</strong>: Conflicting information, need more sources</li>\n</ul>\n<h3 id=\"3-saturation-score\">3. Saturation Score</h3>\n<p>Saturation detects when new pages stop providing novel information.</p>\n<h4 id=\"detection-algorithm\">Detection Algorithm</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Tracks new unique terms per page</span>\n<span class=\"hljs-attr\">new_terms_page_1</span> = <span class=\"hljs-number\">50</span>\n<span class=\"hljs-attr\">new_terms_page_2</span> = <span class=\"hljs-number\">30</span>  <span class=\"hljs-comment\"># 60% of first</span>\n<span class=\"hljs-attr\">new_terms_page_3</span> = <span class=\"hljs-number\">15</span>  <span class=\"hljs-comment\"># 50% of second</span>\n<span class=\"hljs-attr\">new_terms_page_4</span> = <span class=\"hljs-number\">5</span>   <span class=\"hljs-comment\"># 33% of third</span>\n<span class=\"hljs-comment\"># Saturation detected: rapidly diminishing returns</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"configuration\">Configuration</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = AdaptiveConfig(\n    min_gain_threshold=0.1  <span class=\"hljs-comment\"># Stop if &lt;10% new information</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"link-ranking-algorithm\">Link Ranking Algorithm</h2>\n<h3 id=\"expected-information-gain\">Expected Information Gain</h3>\n<p>Each uncrawled link is scored based on:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">ExpectedGain(<span class=\"hljs-built_in\">link</span>) = Relevance × Novelty × Authority\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"1-relevance-scoring\">1. Relevance Scoring</h4>\n<p>Uses BM25 algorithm on link preview text:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">relevance</span> = BM25(link.preview_text, query)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Factors:\n- Term frequency in preview\n- Inverse document frequency\n- Preview length normalization</p>\n<h4 id=\"2-novelty-estimation\">2. Novelty Estimation</h4>\n<p>Measures how different the link appears from already-crawled content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">novelty</span> = <span class=\"hljs-number\">1</span> - max_similarity(preview, knowledge_base)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Prevents crawling duplicate or highly similar pages.</p>\n<h4 id=\"3-authority-calculation\">3. Authority Calculation</h4>\n<p>URL structure and domain analysis:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">authority</span> = f(domain_rank, url_depth, url_structure)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Factors:\n- Domain reputation\n- URL depth (fewer slashes = higher authority)\n- Clean URL structure</p>\n<h2 id=\"domain-specific-configurations\">Domain-Specific Configurations</h2>\n<h3 id=\"technical-documentation\">Technical Documentation</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">tech_doc_config = AdaptiveConfig(\n    confidence_threshold=0.85,\n    max_pages=30,\n    top_k_links=3,\n    min_gain_threshold=0.05  <span class=\"hljs-comment\"># Keep crawling for small gains</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Rationale:\n- High threshold ensures comprehensive coverage\n- Lower gain threshold captures edge cases\n- Moderate link following for depth</p>\n<h3 id=\"news-articles\">News &amp; Articles</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">news_config = AdaptiveConfig(\n    confidence_threshold=0.6,\n    max_pages=10,\n    top_k_links=5,\n    min_gain_threshold=0.15  <span class=\"hljs-comment\"># Stop quickly on repetition</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Rationale:\n- Lower threshold (articles often repeat information)\n- Higher gain threshold (avoid duplicate stories)\n- More links per page (explore different perspectives)</p>\n<h3 id=\"e-commerce\">E-commerce</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">ecommerce_config = AdaptiveConfig(\n    confidence_threshold=0.7,\n    max_pages=20,\n    top_k_links=2,\n    min_gain_threshold=0.1\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Rationale:\n- Balanced threshold for product variations\n- Focused link following (avoid infinite products)\n- Standard gain threshold</p>\n<h3 id=\"research-academic\">Research &amp; Academic</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">research_config = AdaptiveConfig(\n    confidence_threshold=0.9,\n    max_pages=50,\n    top_k_links=4,\n    min_gain_threshold=0.02  <span class=\"hljs-comment\"># Very low - capture citations</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Rationale:\n- Very high threshold for completeness\n- Many pages allowed for thorough research\n- Very low gain threshold to capture references</p>\n<h2 id=\"performance-optimization\">Performance Optimization</h2>\n<h3 id=\"memory-management\">Memory Management</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># For large crawls, use streaming</span>\nconfig = AdaptiveConfig(\n    max_pages=<span class=\"hljs-number\">100</span>,\n    save_state=<span class=\"hljs-literal\">True</span>,\n    state_path=<span class=\"hljs-string\">\"large_crawl.json\"</span>\n)\n\n<span class=\"hljs-comment\"># Periodically clean state</span>\n<span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">len</span>(state.knowledge_base) &gt; <span class=\"hljs-number\">1000</span>:\n    <span class=\"hljs-comment\"># Keep only the top 500 most relevant docs</span>\n    top_content = adaptive.get_relevant_content(top_k=<span class=\"hljs-number\">500</span>)\n    keep_indices = {d[<span class=\"hljs-string\">\"index\"</span>] <span class=\"hljs-keyword\">for</span> d <span class=\"hljs-keyword\">in</span> top_content}\n    state.knowledge_base = [\n        doc <span class=\"hljs-keyword\">for</span> i, doc <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(state.knowledge_base) <span class=\"hljs-keyword\">if</span> i <span class=\"hljs-keyword\">in</span> keep_indices\n    ]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"parallel-processing\">Parallel Processing</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Use multiple start points</span>\n<span class=\"hljs-attr\">start_urls</span> = [\n    <span class=\"hljs-string\">\"https://docs.example.com/intro\"</span>,\n    <span class=\"hljs-string\">\"https://docs.example.com/api\"</span>,\n    <span class=\"hljs-string\">\"https://docs.example.com/guides\"</span>\n]\n\n<span class=\"hljs-comment\"># Crawl in parallel</span>\n<span class=\"hljs-attr\">tasks</span> = [\n    adaptive.digest(url, query)\n    for url in start_urls\n]\n<span class=\"hljs-attr\">results</span> = await asyncio.gather(*tasks)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"debugging-analysis\">Debugging &amp; Analysis</h2>\n<h3 id=\"enable-verbose-logging\">Enable Verbose Logging</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> logging\n\nlogging.basicConfig(level=logging.DEBUG)\nadaptive = AdaptiveCrawler(crawler, config, verbose=<span class=\"hljs-literal\">True</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"analyze-crawl-patterns\">Analyze Crawl Patterns</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># After crawling</span>\nstate = <span class=\"hljs-keyword\">await</span> adaptive.digest(start_url, query)\n\n<span class=\"hljs-comment\"># Analyze link selection</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Link selection order:\"</span>)\n<span class=\"hljs-keyword\">for</span> i, url <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(state.crawl_order):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>. <span class=\"hljs-subst\">{url}</span>\"</span>)\n\n<span class=\"hljs-comment\"># Analyze term discovery</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nTerm discovery rate:\"</span>)\n<span class=\"hljs-keyword\">for</span> i, new_terms <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(state.new_terms_history):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>: <span class=\"hljs-subst\">{new_terms}</span> new terms\"</span>)\n\n<span class=\"hljs-comment\"># Analyze score progression</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nScore progression:\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Coverage: <span class=\"hljs-subst\">{state.metrics[<span class=\"hljs-string\">'coverage_history'</span>]}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Saturation: <span class=\"hljs-subst\">{state.metrics[<span class=\"hljs-string\">'saturation_history'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"export-for-analysis\">Export for Analysis</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Export detailed metrics</span>\n<span class=\"hljs-keyword\">import</span> json\n\nmetrics = {\n    <span class=\"hljs-string\">\"query\"</span>: query,\n    <span class=\"hljs-string\">\"total_pages\"</span>: <span class=\"hljs-built_in\">len</span>(state.crawled_urls),\n    <span class=\"hljs-string\">\"confidence\"</span>: adaptive.confidence,\n    <span class=\"hljs-string\">\"coverage_stats\"</span>: adaptive.coverage_stats,\n    <span class=\"hljs-string\">\"crawl_order\"</span>: state.crawl_order,\n    <span class=\"hljs-string\">\"term_frequencies\"</span>: <span class=\"hljs-built_in\">dict</span>(state.term_frequencies),\n    <span class=\"hljs-string\">\"new_terms_history\"</span>: state.new_terms_history\n}\n\n<span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"crawl_analysis.json\"</span>, <span class=\"hljs-string\">\"w\"</span>) <span class=\"hljs-keyword\">as</span> f:\n    json.dump(metrics, f, indent=<span class=\"hljs-number\">2</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"custom-strategies\">Custom Strategies</h2>\n<h3 id=\"implementing-a-custom-strategy\">Implementing a Custom Strategy</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.adaptive_crawler <span class=\"hljs-keyword\">import</span> CrawlStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">DomainSpecificStrategy</span>(<span class=\"hljs-title class_ inherited__\">CrawlStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">calculate_coverage</span>(<span class=\"hljs-params\">self, state: CrawlState</span>) -&gt; <span class=\"hljs-built_in\">float</span>:\n        <span class=\"hljs-comment\"># Custom coverage calculation</span>\n        <span class=\"hljs-comment\"># e.g., weight certain terms more heavily</span>\n        <span class=\"hljs-keyword\">pass</span>\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">calculate_consistency</span>(<span class=\"hljs-params\">self, state: CrawlState</span>) -&gt; <span class=\"hljs-built_in\">float</span>:\n        <span class=\"hljs-comment\"># Custom consistency logic</span>\n        <span class=\"hljs-comment\"># e.g., domain-specific validation</span>\n        <span class=\"hljs-keyword\">pass</span>\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">rank_links</span>(<span class=\"hljs-params\">self, links: <span class=\"hljs-type\">List</span>[Link], state: CrawlState</span>) -&gt; <span class=\"hljs-type\">List</span>[Link]:\n        <span class=\"hljs-comment\"># Custom link ranking</span>\n        <span class=\"hljs-comment\"># e.g., prioritize specific URL patterns</span>\n        <span class=\"hljs-keyword\">pass</span>\n\n<span class=\"hljs-comment\"># Use custom strategy</span>\nadaptive = AdaptiveCrawler(\n    crawler,\n    config=config,\n    strategy=DomainSpecificStrategy()\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"combining-strategies\">Combining Strategies</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">HybridStrategy</span>(<span class=\"hljs-title class_ inherited__\">CrawlStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self</span>):\n        self.strategies = [\n            TechnicalDocStrategy(),\n            SemanticSimilarityStrategy(),\n            URLPatternStrategy()\n        ]\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">calculate_confidence</span>(<span class=\"hljs-params\">self, state: CrawlState</span>) -&gt; <span class=\"hljs-built_in\">float</span>:\n        <span class=\"hljs-comment\"># Weighted combination of strategies</span>\n        scores = [s.calculate_confidence(state) <span class=\"hljs-keyword\">for</span> s <span class=\"hljs-keyword\">in</span> self.strategies]\n        weights = [<span class=\"hljs-number\">0.5</span>, <span class=\"hljs-number\">0.3</span>, <span class=\"hljs-number\">0.2</span>]\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-built_in\">sum</span>(s * w <span class=\"hljs-keyword\">for</span> s, w <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">zip</span>(scores, weights))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices\">Best Practices</h2>\n<h3 id=\"1-start-conservative\">1. Start Conservative</h3>\n<p>Begin with default settings and adjust based on results:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-meta\"># Start with defaults</span>\nresult = <span class=\"hljs-keyword\">await</span> adaptive.digest(url, query)\n\n<span class=\"hljs-meta\"># Analyze and adjust</span>\n<span class=\"hljs-keyword\">if</span> adaptive.confidence &lt; <span class=\"hljs-number\">0.7</span>:\n    config.max_pages += <span class=\"hljs-number\">10</span>\n    config.confidence_threshold -= <span class=\"hljs-number\">0.1</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-monitor-resource-usage\">2. Monitor Resource Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> psutil\n\n<span class=\"hljs-comment\"># Check memory before large crawls</span>\nmemory_percent = psutil.virtual_memory().percent\n<span class=\"hljs-keyword\">if</span> memory_percent &gt; <span class=\"hljs-number\">80</span>:\n    config.max_pages = <span class=\"hljs-built_in\">min</span>(config.max_pages, <span class=\"hljs-number\">20</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-use-domain-knowledge\">3. Use Domain Knowledge</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># For API documentation</span>\nif <span class=\"hljs-string\">\"api\"</span> in <span class=\"hljs-symbol\">start_url</span><span class=\"hljs-punctuation\">:</span>\n    config.top_k_links <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-number\">2</span>  <span class=\"hljs-comment\"># APIs have clear structure</span>\n\n<span class=\"hljs-comment\"># For blogs</span>\nif <span class=\"hljs-string\">\"blog\"</span> in <span class=\"hljs-symbol\">start_url</span><span class=\"hljs-punctuation\">:</span>\n    config.min_gain_threshold <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-number\">0.2</span>  <span class=\"hljs-comment\"># Avoid similar posts</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"4-validate-results\">4. Validate Results</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Always validate the knowledge base</span>\nrelevant_content = adaptive.get_relevant_content(top_k=<span class=\"hljs-number\">10</span>)\n\n<span class=\"hljs-comment\"># Check coverage</span>\nquery_terms = <span class=\"hljs-built_in\">set</span>(query.lower().split())\ncovered_terms = <span class=\"hljs-built_in\">set</span>()\n\n<span class=\"hljs-keyword\">for</span> doc <span class=\"hljs-keyword\">in</span> relevant_content:\n    content_lower = doc[<span class=\"hljs-string\">'content'</span>].lower()\n    <span class=\"hljs-keyword\">for</span> term <span class=\"hljs-keyword\">in</span> query_terms:\n        <span class=\"hljs-keyword\">if</span> term <span class=\"hljs-keyword\">in</span> content_lower:\n            covered_terms.add(term)\n\ncoverage_ratio = <span class=\"hljs-built_in\">len</span>(covered_terms) / <span class=\"hljs-built_in\">len</span>(query_terms)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Query term coverage: <span class=\"hljs-subst\">{coverage_ratio:<span class=\"hljs-number\">.0</span>%}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"next-steps\">Next Steps</h2>\n<ul>\n<li>Explore <a href=\"../tutorials/custom-adaptive-strategies.md\">Custom Strategy Implementation</a></li>\n<li>Learn about <a href=\"../tutorials/knowledge-base-management.md\">Knowledge Base Management</a></li>\n<li>See <a href=\"../benchmarks/adaptive-performance.md\">Performance Benchmarks</a></li>\n</ul>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/advanced-features/",
    "title": "Overview - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"overview-of-some-important-advanced-features\">Overview of Some Important Advanced Features</h1>\n<p>(Proxy, PDF, Screenshot, SSL, Headers, &amp; Storage State)</p>\n<p>Crawl4AI offers multiple power-user features that go beyond simple crawling. This tutorial covers:</p>\n<p>1. <strong>Proxy Usage</strong><br>\n2. <strong>Capturing PDFs &amp; Screenshots</strong><br>\n3. <strong>Handling SSL Certificates</strong><br>\n4. <strong>Custom Headers</strong><br>\n5. <strong>Session Persistence &amp; Local Storage</strong><br>\n6. <strong>Robots.txt Compliance</strong>  </p>\n<blockquote>\n<p><strong>Prerequisites</strong><br>\n- You have a basic grasp of <a href=\"../../core/simple-crawling/\">AsyncWebCrawler Basics</a><br>\n- You know how to run or configure your Python environment with Playwright installed</p>\n</blockquote>\n<hr>\n<h2 id=\"1-proxy-usage\">1. Proxy Usage</h2>\n<p>If you need to route your crawl traffic through a proxy—whether for IP rotation, geo-testing, or privacy—Crawl4AI supports it via <code>BrowserConfig.proxy_config</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_cfg = BrowserConfig(\n        proxy_config={\n            <span class=\"hljs-string\">\"server\"</span>: <span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>,\n            <span class=\"hljs-string\">\"username\"</span>: <span class=\"hljs-string\">\"myuser\"</span>,\n            <span class=\"hljs-string\">\"password\"</span>: <span class=\"hljs-string\">\"mypass\"</span>,\n        },\n        headless=<span class=\"hljs-literal\">True</span>\n    )\n    crawler_cfg = CrawlerRunConfig(\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.whatismyip.com/\"</span>,\n            config=crawler_cfg\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Page fetched via proxy.\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Page HTML snippet:\"</span>, result.html[:<span class=\"hljs-number\">200</span>])\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR]\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong><br>\n- <strong><code>proxy_config</code></strong> expects a dict with <code>server</code> and optional auth credentials.<br>\n- Many commercial proxies provide an HTTP/HTTPS “gateway” server that you specify in <code>server</code>.<br>\n- If your proxy doesn’t need auth, omit <code>username</code>/<code>password</code>.</p>\n<hr>\n<h2 id=\"2-capturing-pdfs-screenshots\">2. Capturing PDFs &amp; Screenshots</h2>\n<p>Sometimes you need a visual record of a page or a PDF “printout.” Crawl4AI can do both in one pass:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os, asyncio\n<span class=\"hljs-keyword\">from</span> base64 <span class=\"hljs-keyword\">import</span> b64decode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    run_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        pdf=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://en.wikipedia.org/wiki/List_of_common_misconceptions\"</span>,\n            config=run_config\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Screenshot data present: <span class=\"hljs-subst\">{result.screenshot <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"PDF data present: <span class=\"hljs-subst\">{result.pdf <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>}</span>\"</span>)\n\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] Screenshot captured, size: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.screenshot)}</span> bytes\"</span>)\n                <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"wikipedia_screenshot.png\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                    f.write(b64decode(result.screenshot))\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[WARN] Screenshot data is None.\"</span>)\n\n            <span class=\"hljs-keyword\">if</span> result.pdf:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] PDF captured, size: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.pdf)}</span> bytes\"</span>)\n                <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"wikipedia_page.pdf\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                    f.write(result.pdf)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[WARN] PDF data is None.\"</span>)\n\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR]\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Why PDF + Screenshot?</strong><br>\n- Large or complex pages can be slow or error-prone with “traditional” full-page screenshots.<br>\n- Exporting a PDF is more reliable for very long pages. Crawl4AI automatically converts the first PDF page into an image if you request both.  </p>\n<p><strong>Relevant Parameters</strong><br>\n- <strong><code>pdf=True</code></strong>: Exports the current page as a PDF (base64-encoded in <code>result.pdf</code>).<br>\n- <strong><code>screenshot=True</code></strong>: Creates a screenshot (base64-encoded in <code>result.screenshot</code>).<br>\n- <strong><code>scroll_delay</code></strong>: Controls the delay (seconds) between scroll steps when taking a full-page screenshot of a tall page. Defaults to <code>0.2</code>. Increase for pages with slow-loading assets.<br>\n- <strong><code>scan_full_page</code></strong> or advanced hooking can further refine how the crawler captures content.</p>\n<hr>\n<h2 id=\"3-handling-ssl-certificates\">3. Handling SSL Certificates</h2>\n<p>If you need to verify or export a site’s SSL certificate—for compliance, debugging, or data analysis—Crawl4AI can fetch it during the crawl:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio, os\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    tmp_dir = os.path.join(os.getcwd(), <span class=\"hljs-string\">\"tmp\"</span>)\n    os.makedirs(tmp_dir, exist_ok=<span class=\"hljs-literal\">True</span>)\n\n    config = CrawlerRunConfig(\n        fetch_ssl_certificate=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.ssl_certificate:\n            cert = result.ssl_certificate\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nCertificate Information:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Issuer (CN): <span class=\"hljs-subst\">{cert.issuer.get(<span class=\"hljs-string\">'CN'</span>, <span class=\"hljs-string\">''</span>)}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Valid until: <span class=\"hljs-subst\">{cert.valid_until}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Fingerprint: <span class=\"hljs-subst\">{cert.fingerprint}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Export in multiple formats:</span>\n            cert.to_json(os.path.join(tmp_dir, <span class=\"hljs-string\">\"certificate.json\"</span>))\n            cert.to_pem(os.path.join(tmp_dir, <span class=\"hljs-string\">\"certificate.pem\"</span>))\n            cert.to_der(os.path.join(tmp_dir, <span class=\"hljs-string\">\"certificate.der\"</span>))\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nCertificate exported to JSON/PEM/DER in 'tmp' folder.\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR] No certificate or crawl failed.\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong><br>\n- <strong><code>fetch_ssl_certificate=True</code></strong> triggers certificate retrieval.<br>\n- <code>result.ssl_certificate</code> includes methods (<code>to_json</code>, <code>to_pem</code>, <code>to_der</code>) for saving in various formats (handy for server config, Java keystores, etc.).</p>\n<hr>\n<h2 id=\"4-custom-headers\">4. Custom Headers</h2>\n<p>Sometimes you need to set custom headers (e.g., language preferences, authentication tokens, or specialized user-agent strings). You can do this in multiple ways:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Option 1: Set headers at the crawler strategy level</span>\n    crawler1 = AsyncWebCrawler(\n        <span class=\"hljs-comment\"># The underlying strategy can accept headers in its constructor</span>\n        crawler_strategy=<span class=\"hljs-literal\">None</span>  <span class=\"hljs-comment\"># We'll override below for clarity</span>\n    )\n    crawler1.crawler_strategy.update_user_agent(<span class=\"hljs-string\">\"MyCustomUA/1.0\"</span>)\n    crawler1.crawler_strategy.set_custom_headers({\n        <span class=\"hljs-string\">\"Accept-Language\"</span>: <span class=\"hljs-string\">\"fr-FR,fr;q=0.9\"</span>\n    })\n    result1 = <span class=\"hljs-keyword\">await</span> crawler1.arun(<span class=\"hljs-string\">\"https://www.example.com\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Example 1 result success:\"</span>, result1.success)\n\n    <span class=\"hljs-comment\"># Option 2: Pass headers directly to `arun()`</span>\n    crawler2 = AsyncWebCrawler()\n    result2 = <span class=\"hljs-keyword\">await</span> crawler2.arun(\n        url=<span class=\"hljs-string\">\"https://www.example.com\"</span>,\n        headers={<span class=\"hljs-string\">\"Accept-Language\"</span>: <span class=\"hljs-string\">\"es-ES,es;q=0.9\"</span>}\n    )\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Example 2 result success:\"</span>, result2.success)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Notes</strong><br>\n- Some sites may react differently to certain headers (e.g., <code>Accept-Language</code>).<br>\n- If you need advanced user-agent randomization or client hints, see <a href=\"../identity-based-crawling/\">Identity-Based Crawling (Anti-Bot)</a> or use <code>UserAgentGenerator</code>.</p>\n<hr>\n<h2 id=\"5-session-persistence-local-storage\">5. Session Persistence &amp; Local Storage</h2>\n<p>Crawl4AI can preserve cookies and localStorage so you can continue where you left off—ideal for logging into sites or skipping repeated auth flows.</p>\n<h3 id=\"51-storage_state\">5.1 <code>storage_state</code></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    storage_dict = {\n        <span class=\"hljs-string\">\"cookies\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"session\"</span>,\n                <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"abcd1234\"</span>,\n                <span class=\"hljs-string\">\"domain\"</span>: <span class=\"hljs-string\">\"example.com\"</span>,\n                <span class=\"hljs-string\">\"path\"</span>: <span class=\"hljs-string\">\"/\"</span>,\n                <span class=\"hljs-string\">\"expires\"</span>: <span class=\"hljs-number\">1699999999.0</span>,\n                <span class=\"hljs-string\">\"httpOnly\"</span>: <span class=\"hljs-literal\">False</span>,\n                <span class=\"hljs-string\">\"secure\"</span>: <span class=\"hljs-literal\">False</span>,\n                <span class=\"hljs-string\">\"sameSite\"</span>: <span class=\"hljs-string\">\"None\"</span>\n            }\n        ],\n        <span class=\"hljs-string\">\"origins\"</span>: [\n            {\n                <span class=\"hljs-string\">\"origin\"</span>: <span class=\"hljs-string\">\"https://example.com\"</span>,\n                <span class=\"hljs-string\">\"localStorage\"</span>: [\n                    {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"token\"</span>, <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"my_auth_token\"</span>}\n                ]\n            }\n        ]\n    }\n\n    <span class=\"hljs-comment\"># Provide the storage state as a dictionary to start \"already logged in\"</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(\n        headless=<span class=\"hljs-literal\">True</span>,\n        storage_state=storage_dict\n    ) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/protected\"</span>)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Protected page content length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.html))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Failed to crawl protected page\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"52-exporting-reusing-state\">5.2 Exporting &amp; Reusing State</h3>\n<p>You can sign in once, export the browser context, and reuse it later—without re-entering credentials.</p>\n<ul>\n<li><strong><code>await context.storage_state(path=\"my_storage.json\")</code></strong>: Exports cookies, localStorage, etc. to a file.  </li>\n<li>Provide <code>storage_state=\"my_storage.json\"</code> on subsequent runs to skip the login step.</li>\n</ul>\n<p><strong>See</strong>: <a href=\"../session-management/\">Detailed session management tutorial</a> or <a href=\"../identity-based-crawling/\">Explanations → Browser Context &amp; Managed Browser</a> for more advanced scenarios (like multi-step logins, or capturing after interactive pages).</p>\n<hr>\n<h2 id=\"6-robotstxt-compliance\">6. Robots.txt Compliance</h2>\n<p>Crawl4AI supports respecting robots.txt rules with efficient caching:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Enable robots.txt checking in config</span>\n    config = CrawlerRunConfig(\n        check_robots_txt=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Will check and respect robots.txt rules</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            <span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success <span class=\"hljs-keyword\">and</span> result.status_code == <span class=\"hljs-number\">403</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Access denied by robots.txt\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong>\n- Robots.txt files are cached locally for efficiency\n- Cache is stored in <code>~/.crawl4ai/robots/robots_cache.db</code>\n- Cache has a default TTL of 7 days\n- If robots.txt can't be fetched, crawling is allowed\n- Returns 403 status code if URL is disallowed</p>\n<hr>\n<h2 id=\"putting-it-all-together\">Putting It All Together</h2>\n<p>Here’s a snippet that combines multiple “advanced” features (proxy, PDF, screenshot, SSL, custom headers, and session reuse) into one run. Normally, you’d tailor each setting to your project’s needs.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os, asyncio\n<span class=\"hljs-keyword\">from</span> base64 <span class=\"hljs-keyword\">import</span> b64decode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Browser config with proxy + headless</span>\n    browser_cfg = BrowserConfig(\n        proxy_config={\n            <span class=\"hljs-string\">\"server\"</span>: <span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>,\n            <span class=\"hljs-string\">\"username\"</span>: <span class=\"hljs-string\">\"myuser\"</span>,\n            <span class=\"hljs-string\">\"password\"</span>: <span class=\"hljs-string\">\"mypass\"</span>,\n        },\n        headless=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-comment\"># 2. Crawler config with PDF, screenshot, SSL, custom headers, and ignoring caches</span>\n    crawler_cfg = CrawlerRunConfig(\n        pdf=<span class=\"hljs-literal\">True</span>,\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        fetch_ssl_certificate=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.BYPASS,\n        headers={<span class=\"hljs-string\">\"Accept-Language\"</span>: <span class=\"hljs-string\">\"en-US,en;q=0.8\"</span>},\n        storage_state=<span class=\"hljs-string\">\"my_storage.json\"</span>,  <span class=\"hljs-comment\"># Reuse session from a previous sign-in</span>\n        verbose=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-comment\"># 3. Crawl</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url = <span class=\"hljs-string\">\"https://secure.example.com/protected\"</span>, \n            config=crawler_cfg\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Crawled the secure page. Links found:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])))\n\n            <span class=\"hljs-comment\"># Save PDF &amp; screenshot</span>\n            <span class=\"hljs-keyword\">if</span> result.pdf:\n                <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"result.pdf\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                    f.write(b64decode(result.pdf))\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"result.png\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                    f.write(b64decode(result.screenshot))\n\n            <span class=\"hljs-comment\"># Check SSL cert</span>\n            <span class=\"hljs-keyword\">if</span> result.ssl_certificate:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"SSL Issuer CN:\"</span>, result.ssl_certificate.issuer.get(<span class=\"hljs-string\">\"CN\"</span>, <span class=\"hljs-string\">\"\"</span>))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR]\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<hr>\n<h2 id=\"7-anti-bot-features-stealth-mode-undetected-browser\">7. Anti-Bot Features (Stealth Mode &amp; Undetected Browser)</h2>\n<p>Crawl4AI provides two powerful features to bypass bot detection:</p>\n<h3 id=\"71-stealth-mode\">7.1 Stealth Mode</h3>\n<p>Stealth mode uses playwright-stealth to modify browser fingerprints and behaviors. Enable it with a simple flag:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">browser_config <span class=\"hljs-punctuation\">=</span> BrowserConfig<span class=\"hljs-punctuation\">(</span>\n    enable_stealth<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Activates stealth mode</span>\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>When to use</strong>: Sites with basic bot detection (checking navigator.webdriver, plugins, etc.)</p>\n<h3 id=\"72-undetected-browser\">7.2 Undetected Browser</h3>\n<p>For advanced bot detection, use the undetected browser adapter:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> UndetectedAdapter\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_crawler_strategy <span class=\"hljs-keyword\">import</span> AsyncPlaywrightCrawlerStrategy\n\n<span class=\"hljs-comment\"># Create undetected adapter</span>\nadapter = UndetectedAdapter()\nstrategy = AsyncPlaywrightCrawlerStrategy(\n    browser_config=browser_config,\n    browser_adapter=adapter\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(crawler_strategy=strategy, config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Your crawling code</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>When to use</strong>: Sites with sophisticated bot detection (Cloudflare, DataDome, etc.)</p>\n<h3 id=\"73-combining-both\">7.3 Combining Both</h3>\n<p>For maximum evasion, combine stealth mode with undetected browser:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">browser_config <span class=\"hljs-punctuation\">=</span> BrowserConfig<span class=\"hljs-punctuation\">(</span>\n    enable_stealth<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable stealth</span>\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>\n<span class=\"hljs-punctuation\">)</span>\n\nadapter <span class=\"hljs-punctuation\">=</span> UndetectedAdapter<span class=\"hljs-punctuation\">(</span><span class=\"hljs-punctuation\">)</span>  <span class=\"hljs-comment\"># Use undetected browser</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"choosing-the-right-approach\">Choosing the Right Approach</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Detection Level</th>\n<th>Recommended Approach</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No protection</td>\n<td>Regular browser</td>\n</tr>\n<tr>\n<td>Basic checks</td>\n<td>Regular + Stealth mode</td>\n</tr>\n<tr>\n<td>Advanced protection</td>\n<td>Undetected browser</td>\n</tr>\n<tr>\n<td>Maximum evasion</td>\n<td>Undetected + Stealth mode</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Best Practice</strong>: Start with regular browser + stealth mode. Only use undetected browser if needed, as it may be slightly slower.</p>\n<p>See <a href=\"../undetected-browser/\">Undetected Browser Mode</a> for detailed examples.</p>\n<hr>\n<h2 id=\"conclusion-next-steps\">Conclusion &amp; Next Steps</h2>\n<p>You've now explored several <strong>advanced</strong> features:</p>\n<ul>\n<li><strong>Proxy Usage</strong>  </li>\n<li><strong>PDF &amp; Screenshot</strong> capturing for large or critical pages  </li>\n<li><strong>SSL Certificate</strong> retrieval &amp; exporting  </li>\n<li><strong>Custom Headers</strong> for language or specialized requests  </li>\n<li><strong>Session Persistence</strong> via storage state</li>\n<li><strong>Robots.txt Compliance</strong></li>\n<li><strong>Anti-Bot Features</strong> (Stealth Mode &amp; Undetected Browser)</li>\n</ul>\n<p>With these power tools, you can build robust scraping workflows that mimic real user behavior, handle secure sites, capture detailed snapshots, manage sessions across multiple runs, and bypass bot detection—streamlining your entire data collection pipeline.</p>\n<p><strong>Note</strong>: In future versions, we may enable stealth mode and undetected browser by default. For now, users should explicitly enable these features when needed.</p>\n<p><strong>Last Updated</strong>: 2025-01-17</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/anti-bot-and-fallback/",
    "title": "Anti-Bot & Fallback - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"anti-bot-detection-fallback\">Anti-Bot Detection &amp; Fallback</h1>\n<p>When crawling sites protected by anti-bot systems (Akamai, Cloudflare, PerimeterX, DataDome, Imperva, etc.), requests often get blocked with CAPTCHAs, 403 responses, or empty pages. Crawl4AI provides a layered retry and fallback system that automatically detects blocking and escalates through multiple strategies until content is retrieved.</p>\n<h2 id=\"how-detection-works\">How Detection Works</h2>\n<p>After each crawl attempt, Crawl4AI inspects the HTTP status code and HTML content for known anti-bot signals:</p>\n<ul>\n<li><strong>HTTP 403/429</strong> with short or empty response bodies</li>\n<li><strong>Challenge pages</strong> — Cloudflare \"Just a moment\", Akamai \"Access Denied\", PerimeterX block pages</li>\n<li><strong>CAPTCHA injection</strong> — reCAPTCHA, hCaptcha, or vendor-specific challenges on otherwise empty pages</li>\n<li><strong>Firewall blocks</strong> — Imperva/Incapsula resource iframes, Sucuri firewall pages, Cloudflare error codes</li>\n</ul>\n<p>Detection uses structural HTML markers (specific element IDs, script sources, form actions) rather than generic keywords to minimize false positives. A normal page that happens to mention \"CAPTCHA\" or \"Cloudflare\" in its content will not be flagged.</p>\n<p>When all attempts fail and blocking is still detected, the result is returned with <code>success=False</code> and <code>error_message</code> describing the block reason.</p>\n<h2 id=\"configuration-options\">Configuration Options</h2>\n<p>All anti-bot retry options live on <code>CrawlerRunConfig</code>:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Type</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>proxy_config</code></td>\n<td><code>ProxyConfig</code>, <code>list[ProxyConfig]</code>, or <code>None</code></td>\n<td><code>None</code></td>\n<td>Single proxy or ordered list of proxies to try. Each retry round iterates through the full list. Use <code>\"direct\"</code> or <code>ProxyConfig.DIRECT</code> in a list to explicitly try without a proxy.</td>\n</tr>\n<tr>\n<td><code>max_retries</code></td>\n<td><code>int</code></td>\n<td><code>0</code></td>\n<td>Number of retry rounds when blocking is detected. <code>0</code> = no retries.</td>\n</tr>\n<tr>\n<td><code>fallback_fetch_function</code></td>\n<td><code>async (str) -&gt; str</code></td>\n<td><code>None</code></td>\n<td>Async function called as last resort. Takes URL, returns raw HTML.</td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"escalation-chain\">Escalation Chain</h2>\n<p>Each retry round tries every proxy in <code>proxy_config</code> in order. If all rounds are exhausted and the page is still blocked, the fallback fetch function is called as a last resort.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">For each <span class=\"hljs-built_in\">round</span> (<span class=\"hljs-number\">1</span> + max_retries rounds):\n    <span class=\"hljs-number\">1.</span> Try proxy_config[<span class=\"hljs-number\">0</span>] (<span class=\"hljs-keyword\">or</span> direct <span class=\"hljs-keyword\">if</span> proxy_config <span class=\"hljs-keyword\">is</span> <span class=\"hljs-literal\">None</span>)\n    <span class=\"hljs-number\">2.</span> If blocked → <span class=\"hljs-keyword\">try</span> proxy_config[<span class=\"hljs-number\">1</span>]\n    <span class=\"hljs-number\">3.</span> If blocked → <span class=\"hljs-keyword\">try</span> proxy_config[<span class=\"hljs-number\">2</span>]\n    <span class=\"hljs-number\">4.</span> ... <span class=\"hljs-keyword\">continue</span> through <span class=\"hljs-built_in\">all</span> proxies\n    <span class=\"hljs-number\">5.</span> If <span class=\"hljs-built_in\">any</span> attempt succeeds → done\n\nIf <span class=\"hljs-built_in\">all</span> rounds exhausted <span class=\"hljs-keyword\">and</span> still blocked:\n    <span class=\"hljs-number\">6.</span> Call fallback_fetch_function(url) → process returned HTML\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Worst-case attempts before the fetch function: <code>(1 + max_retries) x len(proxy_config)</code></p>\n<h2 id=\"crawl-stats\">Crawl Stats</h2>\n<p>Every crawl result includes a <code>crawl_stats</code> dict with detailed attempt tracking:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">result.crawl_stats = {\n    <span class=\"hljs-string\">\"attempts\"</span>: <span class=\"hljs-number\">3</span>,                    <span class=\"hljs-comment\"># total browser attempts made</span>\n    <span class=\"hljs-string\">\"retries\"</span>: <span class=\"hljs-number\">1</span>,                     <span class=\"hljs-comment\"># retry rounds used (0 = succeeded first round)</span>\n    <span class=\"hljs-string\">\"proxies_used\"</span>: [                 <span class=\"hljs-comment\"># ordered list of every attempt</span>\n        {<span class=\"hljs-string\">\"proxy\"</span>: <span class=\"hljs-literal\">None</span>,               <span class=\"hljs-string\">\"status_code\"</span>: <span class=\"hljs-number\">403</span>, <span class=\"hljs-string\">\"blocked\"</span>: <span class=\"hljs-literal\">True</span>,  <span class=\"hljs-string\">\"reason\"</span>: <span class=\"hljs-string\">\"Akamai block (Reference #)\"</span>},\n        {<span class=\"hljs-string\">\"proxy\"</span>: <span class=\"hljs-string\">\"proxy.io:8080\"</span>,    <span class=\"hljs-string\">\"status_code\"</span>: <span class=\"hljs-number\">403</span>, <span class=\"hljs-string\">\"blocked\"</span>: <span class=\"hljs-literal\">True</span>,  <span class=\"hljs-string\">\"reason\"</span>: <span class=\"hljs-string\">\"Akamai block (Reference #)\"</span>},\n        {<span class=\"hljs-string\">\"proxy\"</span>: <span class=\"hljs-string\">\"premium.io:9090\"</span>,  <span class=\"hljs-string\">\"status_code\"</span>: <span class=\"hljs-number\">200</span>, <span class=\"hljs-string\">\"blocked\"</span>: <span class=\"hljs-literal\">False</span>, <span class=\"hljs-string\">\"reason\"</span>: <span class=\"hljs-string\">\"\"</span>},\n    ],\n    <span class=\"hljs-string\">\"fallback_fetch_used\"</span>: <span class=\"hljs-literal\">False</span>,     <span class=\"hljs-comment\"># whether fallback_fetch_function was called</span>\n    <span class=\"hljs-string\">\"resolved_by\"</span>: <span class=\"hljs-string\">\"proxy\"</span>,           <span class=\"hljs-comment\"># \"direct\" | \"proxy\" | \"fallback_fetch\" | null (all failed)</span>\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"usage-examples\">Usage Examples</h2>\n<h3 id=\"simple-retry-no-proxy\">Simple Retry (No Proxy)</h3>\n<p>Retry the crawl up to 3 times when blocking is detected. Useful when blocks are intermittent or IP-based.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        config=CrawlerRunConfig(max_retries=<span class=\"hljs-number\">3</span>),\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"single-proxy\">Single Proxy</h3>\n<p>Pass a single <code>ProxyConfig</code> — it's used on every attempt. Same behavior as always.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-java\">from crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">ProxyConfig</span>\n\n<span class=\"hljs-variable\">config</span> <span class=\"hljs-operator\">=</span> CrawlerRunConfig(\n    max_retries=<span class=\"hljs-number\">2</span>,\n    proxy_config=ProxyConfig(\n        server=<span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>,\n        username=<span class=\"hljs-string\">\"user\"</span>,\n        password=<span class=\"hljs-string\">\"pass\"</span>,\n    ),\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"direct-first-then-proxies\">Direct-First, Then Proxies</h3>\n<p>Try without a proxy first, then escalate to proxies if blocked. Use <code>ProxyConfig.DIRECT</code> (or the string <code>\"direct\"</code>) in the list to represent a no-proxy attempt.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    max_retries=1,\n    proxy_config=[\n        ProxyConfig.DIRECT,  <span class=\"hljs-comment\"># Try without proxy first</span>\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://datacenter-proxy.example.com:8080\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://residential-proxy.example.com:9090\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n    ],\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>With this setup, each round tries direct first, then datacenter, then residential. With <code>max_retries=1</code>, worst case is 2 rounds x 3 steps = 6 attempts.</p>\n<h3 id=\"proxy-list-escalation\">Proxy List (Escalation)</h3>\n<p>Pass a list of proxies. They're tried in order — first one that works wins. Within each retry round, the entire list is tried again.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-built_in\">config</span> = CrawlerRunConfig(\n    max_retries=<span class=\"hljs-number\">1</span>,\n    proxy_config=[\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://datacenter-proxy.example.com:8080\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://residential-proxy.example.com:9090\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n    ],\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>With this setup, each round tries the datacenter proxy first, then the residential proxy. With <code>max_retries=1</code>, worst case is 2 rounds x 2 proxies = 4 attempts.</p>\n<h3 id=\"fallback-fetch-function\">Fallback Fetch Function</h3>\n<p>When all browser-based attempts fail, call a custom async function as a last resort. This function receives the URL and must return raw HTML as a string. The returned HTML is processed through the normal pipeline (markdown generation, extraction, etc.).</p>\n<p>This is useful when you have access to a scraping API, a pre-fetched cache, or any other source of HTML.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> aiohttp\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">my_scraping_api</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span></span>) -&gt; <span class=\"hljs-built_in\">str</span>:\n    <span class=\"hljs-string\">\"\"\"Fetch HTML via an external scraping API.\"\"\"</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> aiohttp.ClientSession() <span class=\"hljs-keyword\">as</span> session:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> session.get(\n            <span class=\"hljs-string\">\"https://api.my-scraping-service.com/fetch\"</span>,\n            params={<span class=\"hljs-string\">\"url\"</span>: url, <span class=\"hljs-string\">\"format\"</span>: <span class=\"hljs-string\">\"html\"</span>},\n            headers={<span class=\"hljs-string\">\"Authorization\"</span>: <span class=\"hljs-string\">\"Bearer MY_TOKEN\"</span>},\n        ) <span class=\"hljs-keyword\">as</span> resp:\n            <span class=\"hljs-keyword\">if</span> resp.status == <span class=\"hljs-number\">200</span>:\n                <span class=\"hljs-keyword\">return</span> <span class=\"hljs-keyword\">await</span> resp.text()\n            <span class=\"hljs-keyword\">raise</span> RuntimeError(<span class=\"hljs-string\">f\"API error: <span class=\"hljs-subst\">{resp.status}</span>\"</span>)\n\nconfig = CrawlerRunConfig(\n    max_retries=<span class=\"hljs-number\">1</span>,\n    fallback_fetch_function=my_scraping_api,\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The function can do anything — call an API, read from a database, return cached HTML, or make a simple HTTP request with a different library. Crawl4AI does not care how the HTML is obtained.</p>\n<h3 id=\"full-escalation-all-features-combined\">Full Escalation (All Features Combined)</h3>\n<p>This example combines every layer: stealth mode, a list of proxies tried in order, retries, and a final fetch function.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> aiohttp\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig, ProxyConfig\n\n<span class=\"hljs-comment\"># Last-resort: fetch HTML via an external service</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">external_fetch</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span></span>) -&gt; <span class=\"hljs-built_in\">str</span>:\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> aiohttp.ClientSession() <span class=\"hljs-keyword\">as</span> session:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> session.post(\n            <span class=\"hljs-string\">\"https://api.my-service.com/scrape\"</span>,\n            json={<span class=\"hljs-string\">\"url\"</span>: url, <span class=\"hljs-string\">\"render_js\"</span>: <span class=\"hljs-literal\">True</span>},\n            headers={<span class=\"hljs-string\">\"Authorization\"</span>: <span class=\"hljs-string\">\"Bearer MY_TOKEN\"</span>},\n        ) <span class=\"hljs-keyword\">as</span> resp:\n            <span class=\"hljs-keyword\">return</span> <span class=\"hljs-keyword\">await</span> resp.text()\n\nbrowser_config = BrowserConfig(\n    headless=<span class=\"hljs-literal\">True</span>,\n    enable_stealth=<span class=\"hljs-literal\">True</span>,\n)\n\ncrawl_config = CrawlerRunConfig(\n    magic=<span class=\"hljs-literal\">True</span>,\n    wait_until=<span class=\"hljs-string\">\"load\"</span>,\n    max_retries=<span class=\"hljs-number\">2</span>,\n\n    <span class=\"hljs-comment\"># Proxies tried in order — cheapest first</span>\n    proxy_config=[\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://datacenter-proxy.example.com:8080\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n        ProxyConfig(\n            server=<span class=\"hljs-string\">\"http://residential-proxy.example.com:9090\"</span>,\n            username=<span class=\"hljs-string\">\"user\"</span>,\n            password=<span class=\"hljs-string\">\"pass\"</span>,\n        ),\n    ],\n\n    <span class=\"hljs-comment\"># Last resort — called after all retries and proxies are exhausted</span>\n    fallback_fetch_function=external_fetch,\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://protected-site.com/products\"</span>,\n        config=crawl_config,\n    )\n\n    <span class=\"hljs-keyword\">if</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Got <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown)}</span> chars of markdown\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Resolved by: <span class=\"hljs-subst\">{result.crawl_stats[<span class=\"hljs-string\">'resolved_by'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Attempts: <span class=\"hljs-subst\">{result.crawl_stats[<span class=\"hljs-string\">'attempts'</span>]}</span>\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"All attempts failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What happens step by step:</strong></p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Round</th>\n<th>Attempt</th>\n<th>What runs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>Datacenter proxy — blocked</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2</td>\n<td>Residential proxy — blocked</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1</td>\n<td>Datacenter proxy — blocked</td>\n</tr>\n<tr>\n<td>2</td>\n<td>2</td>\n<td>Residential proxy — blocked</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1</td>\n<td>Datacenter proxy — blocked</td>\n</tr>\n<tr>\n<td>3</td>\n<td>2</td>\n<td>Residential proxy — blocked</td>\n</tr>\n<tr>\n<td>-</td>\n<td>-</td>\n<td><code>external_fetch(url)</code> called — returns HTML</td>\n</tr>\n</tbody>\n</table>\n<p>That's up to 6 browser attempts + 1 function call before giving up.</p>\n<h2 id=\"tips\">Tips</h2>\n<ul>\n<li><strong>Start with <code>max_retries=0</code></strong> and a <code>fallback_fetch_function</code> if you just want a safety net without burning time on retries.</li>\n<li><strong>Order proxies cheapest-first</strong> — datacenter proxies before residential, residential before premium.</li>\n<li><strong>Combine with stealth mode</strong> — <code>BrowserConfig(enable_stealth=True)</code> and <code>CrawlerRunConfig(magic=True)</code> reduce the chance of being blocked in the first place.</li>\n<li><strong><code>wait_until=\"load\"</code></strong> is important for anti-bot sites — the default <code>domcontentloaded</code> can return before the anti-bot sensor finishes.</li>\n<li><strong>Check <code>crawl_stats</code></strong> to understand what happened — how many attempts, which proxy worked, whether the fallback function was needed.</li>\n</ul>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../proxy-security/\">Proxy &amp; Security</a> — Proxy setup, authentication, and rotation</li>\n<li><a href=\"../undetected-browser/\">Undetected Browser</a> — Stealth mode and browser fingerprint evasion</li>\n<li><a href=\"../session-management/\">Session Management</a> — Maintaining sessions across requests</li>\n</ul>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/crawl-dispatcher/",
    "title": "Crawl Dispatcher - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl-dispatcher\">Crawl Dispatcher</h1>\n<p>We’re excited to announce a <strong>Crawl Dispatcher</strong> module that can handle <strong>thousands</strong> of crawling tasks simultaneously. By efficiently managing system resources (memory, CPU, network), this dispatcher ensures high-performance data extraction at scale. It also provides <strong>real-time monitoring</strong> of each crawler’s status, memory usage, and overall progress.</p>\n<p>Stay tuned—this feature is <strong>coming soon</strong> in an upcoming release of Crawl4AI! For the latest news, keep an eye on our changelogs and follow <a href=\"https://twitter.com/unclecode\">@unclecode</a> on X.</p>\n<p>Below is a <strong>sample</strong> of how the dispatcher’s performance monitor might look in action:</p>\n<p><img alt=\"Crawl Dispatcher Performance Monitor\" src=\"../../assets/images/dispatcher.png\"></p>\n<p>We can’t wait to bring you this streamlined, <strong>scalable</strong> approach to multi-URL crawling—<strong>watch this space</strong> for updates!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/file-downloading/",
    "title": "File Downloading - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"download-handling-in-crawl4ai\">Download Handling in Crawl4AI</h1>\n<p>This guide explains how to use Crawl4AI to handle file downloads during crawling. You'll learn how to trigger downloads, specify download locations, and access downloaded files.</p>\n<h2 id=\"enabling-downloads\">Enabling Downloads</h2>\n<p>To enable downloads, set the <code>accept_downloads</code> parameter in the <code>BrowserConfig</code> object and pass it to the crawler.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = BrowserConfig(accept_downloads=<span class=\"hljs-literal\">True</span>)  <span class=\"hljs-comment\"># Enable downloads globally</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=config) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># ... your crawling logic ...</span>\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"specifying-download-location\">Specifying Download Location</h2>\n<p>Specify the download directory using the <code>downloads_path</code> attribute in the <code>BrowserConfig</code> object. If not provided, Crawl4AI defaults to creating a \"downloads\" directory inside the <code>.crawl4ai</code> folder in your home directory.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig\n<span class=\"hljs-keyword\">import</span> os\n\ndownloads_path = os.path.join(os.getcwd(), <span class=\"hljs-string\">\"my_downloads\"</span>)  <span class=\"hljs-comment\"># Custom download path</span>\nos.makedirs(downloads_path, exist_ok=<span class=\"hljs-literal\">True</span>)\n\nconfig = BrowserConfig(accept_downloads=<span class=\"hljs-literal\">True</span>, downloads_path=downloads_path)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>)\n        <span class=\"hljs-comment\"># ...</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"triggering-downloads\">Triggering Downloads</h2>\n<p>Downloads are typically triggered by user interactions on a web page, such as clicking a download button. Use <code>js_code</code> in <code>CrawlerRunConfig</code> to simulate these actions and <code>wait_for</code> to allow sufficient time for downloads to start.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\nconfig = CrawlerRunConfig(\n    js_code=<span class=\"hljs-string\">\"\"\"\n        const downloadLink = document.querySelector('a[href$=\".exe\"]');\n        if (downloadLink) {\n            downloadLink.click();\n        }\n    \"\"\"</span>,\n    wait_for=<span class=\"hljs-number\">5</span>  <span class=\"hljs-comment\"># Wait 5 seconds for the download to start</span>\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://www.python.org/downloads/\"</span>, config=config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"accessing-downloaded-files\">Accessing Downloaded Files</h2>\n<p>The <code>downloaded_files</code> attribute of the <code>CrawlResult</code> object contains paths to downloaded files.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-keyword\">if</span> result.downloaded_files:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Downloaded files:\"</span>)\n    <span class=\"hljs-keyword\">for</span> file_path <span class=\"hljs-keyword\">in</span> result.downloaded_files:\n        <span class=\"hljs-built_in\">print</span>(f<span class=\"hljs-string\">\"- {file_path}\"</span>)\n        file_size = <span class=\"hljs-built_in\">os</span>.<span class=\"hljs-built_in\">path</span>.getsize(file_path)\n        <span class=\"hljs-built_in\">print</span>(f<span class=\"hljs-string\">\"- File size: {file_size} bytes\"</span>)\n<span class=\"hljs-keyword\">else</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"No files downloaded.\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"example-downloading-multiple-files\">Example: Downloading Multiple Files</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n<span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">from</span> pathlib <span class=\"hljs-keyword\">import</span> Path\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">download_multiple_files</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span>, download_path: <span class=\"hljs-built_in\">str</span></span>):\n    config = BrowserConfig(accept_downloads=<span class=\"hljs-literal\">True</span>, downloads_path=download_path)\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=config) <span class=\"hljs-keyword\">as</span> crawler:\n        run_config = CrawlerRunConfig(\n            js_code=<span class=\"hljs-string\">\"\"\"\n                const downloadLinks = document.querySelectorAll('a[download]');\n                for (const link of downloadLinks) {\n                    link.click();\n                    // Delay between clicks\n                    await new Promise(r =&gt; setTimeout(r, 2000));  \n                }\n            \"\"\"</span>,\n            wait_for=<span class=\"hljs-number\">10</span>  <span class=\"hljs-comment\"># Wait for all downloads to start</span>\n        )\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=run_config)\n\n        <span class=\"hljs-keyword\">if</span> result.downloaded_files:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Downloaded files:\"</span>)\n            <span class=\"hljs-keyword\">for</span> file <span class=\"hljs-keyword\">in</span> result.downloaded_files:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"- <span class=\"hljs-subst\">{file}</span>\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"No files downloaded.\"</span>)\n\n<span class=\"hljs-comment\"># Usage</span>\ndownload_path = os.path.join(Path.home(), <span class=\"hljs-string\">\".crawl4ai\"</span>, <span class=\"hljs-string\">\"downloads\"</span>)\nos.makedirs(download_path, exist_ok=<span class=\"hljs-literal\">True</span>)\n\nasyncio.run(download_multiple_files(<span class=\"hljs-string\">\"https://www.python.org/downloads/windows/\"</span>, download_path))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"important-considerations\">Important Considerations</h2>\n<ul>\n<li><strong>Browser Context:</strong> Downloads are managed within the browser context. Ensure <code>js_code</code> correctly targets the download triggers on the webpage.</li>\n<li><strong>Timing:</strong> Use <code>wait_for</code> in <code>CrawlerRunConfig</code> to manage download timing.</li>\n<li><strong>Error Handling:</strong> Handle errors to manage failed downloads or incorrect paths gracefully.</li>\n<li><strong>Security:</strong> Scan downloaded files for potential security threats before use.</li>\n</ul>\n<p>This revised guide ensures consistency with the <code>Crawl4AI</code> codebase by using <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> for all download-related configurations. Let me know if further adjustments are needed!</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/hooks-auth/",
    "title": "Hooks & Auth - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"hooks-auth-in-asyncwebcrawler\">Hooks &amp; Auth in AsyncWebCrawler</h1>\n<p>Crawl4AI’s <strong>hooks</strong> let you customize the crawler at specific points in the pipeline:</p>\n<p>1. <strong><code>on_browser_created</code></strong> – After browser creation.<br>\n2. <strong><code>on_page_context_created</code></strong> – After a new context &amp; page are created.<br>\n3. <strong><code>before_goto</code></strong> – Just before navigating to a page.<br>\n4. <strong><code>after_goto</code></strong> – Right after navigation completes.<br>\n5. <strong><code>on_user_agent_updated</code></strong> – Whenever the user agent changes.<br>\n6. <strong><code>on_execution_started</code></strong> – Once custom JavaScript execution begins.<br>\n7. <strong><code>before_retrieve_html</code></strong> – Just before the crawler retrieves final HTML.<br>\n8. <strong><code>before_return_html</code></strong> – Right before returning the HTML content.</p>\n<p><strong>Important</strong>: Avoid heavy tasks in <code>on_browser_created</code> since you don’t yet have a page context. If you need to <em>log in</em>, do so in <strong><code>on_page_context_created</code></strong>.</p>\n<blockquote>\n<p>note \"Important Hook Usage Warning\"\n    <strong>Avoid Misusing Hooks</strong>: Do not manipulate page objects in the wrong hook or at the wrong time, as it can crash the pipeline or produce incorrect results. A common mistake is attempting to handle authentication prematurely—such as creating or closing pages in <code>on_browser_created</code>. </p>\n<p><strong>Use the Right Hook for Auth</strong>: If you need to log in or set tokens, use <code>on_page_context_created</code>. This ensures you have a valid page/context to work with, without disrupting the main crawling flow.</p>\n<p><strong>Identity-Based Crawling</strong>: For robust auth, consider identity-based crawling (or passing a session ID) to preserve state. Run your initial login steps in a separate, well-defined process, then feed that session to your main crawl—rather than shoehorning complex authentication into early hooks. Check out <a href=\"../identity-based-crawling/\">Identity-Based Crawling</a> for more details.</p>\n<p><strong>Be Cautious</strong>: Overwriting or removing elements in the wrong hook can compromise the final crawl. Keep hooks focused on smaller tasks (like route filters, custom headers), and let your main logic (crawling, data extraction) proceed normally.</p>\n</blockquote>\n<p>Below is an example demonstration.</p>\n<hr>\n<h2 id=\"example-using-hooks-in-asyncwebcrawler\">Example: Using Hooks in AsyncWebCrawler</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> playwright.async_api <span class=\"hljs-keyword\">import</span> Page, BrowserContext\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"🔗 Hooks Example: Demonstrating recommended usage\"</span>)\n\n    <span class=\"hljs-comment\"># 1) Configure the browser</span>\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># 2) Configure the crawler run</span>\n    crawler_run_config = CrawlerRunConfig(\n        js_code=<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        wait_for=<span class=\"hljs-string\">\"body\"</span>,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-comment\"># 3) Create the crawler instance</span>\n    crawler = AsyncWebCrawler(config=browser_config)\n\n    <span class=\"hljs-comment\">#</span>\n    <span class=\"hljs-comment\"># Define Hook Functions</span>\n    <span class=\"hljs-comment\">#</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_browser_created</span>(<span class=\"hljs-params\">browser, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called once the browser instance is created (but no pages or contexts yet)</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_browser_created - Browser created successfully!\"</span>)\n        <span class=\"hljs-comment\"># Typically, do minimal setup here if needed</span>\n        <span class=\"hljs-keyword\">return</span> browser\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_page_context_created</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called right after a new page + context are created (ideal for auth or route config).</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_page_context_created - Setting up page &amp; context.\"</span>)\n\n        <span class=\"hljs-comment\"># Example 1: Route filtering (e.g., block images)</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">route_filter</span>(<span class=\"hljs-params\">route</span>):\n            <span class=\"hljs-keyword\">if</span> route.request.resource_type == <span class=\"hljs-string\">\"image\"</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] Blocking image request: <span class=\"hljs-subst\">{route.request.url}</span>\"</span>)\n                <span class=\"hljs-keyword\">await</span> route.abort()\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-keyword\">await</span> route.continue_()\n\n        <span class=\"hljs-keyword\">await</span> context.route(<span class=\"hljs-string\">\"**\"</span>, route_filter)\n\n        <span class=\"hljs-comment\"># Example 2: (Optional) Simulate a login scenario</span>\n        <span class=\"hljs-comment\"># (We do NOT create or close pages here, just do quick steps if needed)</span>\n        <span class=\"hljs-comment\"># e.g., await page.goto(\"https://example.com/login\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.fill(\"input[name='username']\", \"testuser\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.fill(\"input[name='password']\", \"password123\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.click(\"button[type='submit']\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.wait_for_selector(\"#welcome\")</span>\n        <span class=\"hljs-comment\"># e.g., await context.add_cookies([...])</span>\n        <span class=\"hljs-comment\"># Then continue</span>\n\n        <span class=\"hljs-comment\"># Example 3: Adjust the viewport</span>\n        <span class=\"hljs-keyword\">await</span> page.set_viewport_size({<span class=\"hljs-string\">\"width\"</span>: <span class=\"hljs-number\">1080</span>, <span class=\"hljs-string\">\"height\"</span>: <span class=\"hljs-number\">600</span>})\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_goto</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, url: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called before navigating to each URL.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] before_goto - About to navigate: <span class=\"hljs-subst\">{url}</span>\"</span>)\n        <span class=\"hljs-comment\"># e.g., inject custom headers</span>\n        <span class=\"hljs-keyword\">await</span> page.set_extra_http_headers({\n            <span class=\"hljs-string\">\"Custom-Header\"</span>: <span class=\"hljs-string\">\"my-value\"</span>\n        })\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">after_goto</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, \n        url: <span class=\"hljs-built_in\">str</span>, response, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called after navigation completes.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] after_goto - Successfully loaded: <span class=\"hljs-subst\">{url}</span>\"</span>)\n        <span class=\"hljs-comment\"># e.g., wait for a certain element if we want to verify</span>\n        <span class=\"hljs-keyword\">try</span>:\n            <span class=\"hljs-keyword\">await</span> page.wait_for_selector(<span class=\"hljs-string\">'.content'</span>, timeout=<span class=\"hljs-number\">1000</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] Found .content element!\"</span>)\n        <span class=\"hljs-keyword\">except</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] .content not found, continuing anyway.\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_user_agent_updated</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, \n        user_agent: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called whenever the user agent updates.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] on_user_agent_updated - New user agent: <span class=\"hljs-subst\">{user_agent}</span>\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_execution_started</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called after custom JavaScript execution begins.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_execution_started - JS code is running!\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_retrieve_html</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called before final HTML retrieval.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] before_retrieve_html - We can do final actions\"</span>)\n        <span class=\"hljs-comment\"># Example: Scroll again</span>\n        <span class=\"hljs-keyword\">await</span> page.evaluate(<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_return_html</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, html: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called just before returning the HTML in the result.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] before_return_html - HTML length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(html)}</span>\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-comment\">#</span>\n    <span class=\"hljs-comment\"># Attach Hooks</span>\n    <span class=\"hljs-comment\">#</span>\n\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"on_browser_created\"</span>, on_browser_created)\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_page_context_created\"</span>, on_page_context_created\n    )\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"before_goto\"</span>, before_goto)\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"after_goto\"</span>, after_goto)\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_user_agent_updated\"</span>, on_user_agent_updated\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_execution_started\"</span>, on_execution_started\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"before_retrieve_html\"</span>, before_retrieve_html\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"before_return_html\"</span>, before_return_html\n    )\n\n    <span class=\"hljs-keyword\">await</span> crawler.start()\n\n    <span class=\"hljs-comment\"># 4) Run the crawler on an example page</span>\n    url = <span class=\"hljs-string\">\"https://example.com\"</span>\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config=crawler_run_config)\n\n    <span class=\"hljs-keyword\">if</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nCrawled URL:\"</span>, result.url)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.html))\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n    <span class=\"hljs-keyword\">await</span> crawler.close()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"hook-lifecycle-summary\">Hook Lifecycle Summary</h2>\n<p>1. <strong><code>on_browser_created</code></strong>:<br>\n   - Browser is up, but <strong>no</strong> pages or contexts yet.<br>\n   - Light setup only—don’t try to open or close pages here (that belongs in <code>on_page_context_created</code>).</p>\n<p>2. <strong><code>on_page_context_created</code></strong>:<br>\n   - Perfect for advanced <strong>auth</strong> or route blocking.<br>\n   - You have a <strong>page</strong> + <strong>context</strong> ready but haven’t navigated to the target URL yet.</p>\n<p>3. <strong><code>before_goto</code></strong>:<br>\n   - Right before navigation. Typically used for setting <strong>custom headers</strong> or logging the target URL.</p>\n<p>4. <strong><code>after_goto</code></strong>:<br>\n   - After page navigation is done. Good place for verifying content or waiting on essential elements. </p>\n<p>5. <strong><code>on_user_agent_updated</code></strong>:<br>\n   - Whenever the user agent changes (for stealth or different UA modes).</p>\n<p>6. <strong><code>on_execution_started</code></strong>:<br>\n   - If you set <code>js_code</code> or run custom scripts, this runs once your JS is about to start.</p>\n<p>7. <strong><code>before_retrieve_html</code></strong>:<br>\n   - Just before the final HTML snapshot is taken. Often you do a final scroll or lazy-load triggers here.</p>\n<p>8. <strong><code>before_return_html</code></strong>:<br>\n   - The last hook before returning HTML to the <code>CrawlResult</code>. Good for logging HTML length or minor modifications.</p>\n<hr>\n<h2 id=\"when-to-handle-authentication\">When to Handle Authentication</h2>\n<p><strong>Recommended</strong>: Use <strong><code>on_page_context_created</code></strong> if you need to:</p>\n<ul>\n<li>Navigate to a login page or fill forms</li>\n<li>Set cookies or localStorage tokens</li>\n<li>Block resource routes to avoid ads</li>\n</ul>\n<p>This ensures the newly created context is under your control <strong>before</strong> <code>arun()</code> navigates to the main URL.</p>\n<hr>\n<h2 id=\"additional-considerations\">Additional Considerations</h2>\n<ul>\n<li><strong>Session Management</strong>: If you want multiple <code>arun()</code> calls to reuse a single session, pass <code>session_id=</code> in your <code>CrawlerRunConfig</code>. Hooks remain the same.  </li>\n<li><strong>Performance</strong>: Hooks can slow down crawling if they do heavy tasks. Keep them concise.  </li>\n<li><strong>Error Handling</strong>: If a hook fails, the overall crawl might fail. Catch exceptions or handle them gracefully.  </li>\n<li><strong>Concurrency</strong>: If you run <code>arun_many()</code>, each URL triggers these hooks in parallel. Ensure your hooks are thread/async-safe.</li>\n</ul>\n<hr>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Hooks provide <strong>fine-grained</strong> control over:</p>\n<ul>\n<li><strong>Browser</strong> creation (light tasks only)</li>\n<li><strong>Page</strong> and <strong>context</strong> creation (auth, route blocking)</li>\n<li><strong>Navigation</strong> phases</li>\n<li><strong>Final HTML</strong> retrieval</li>\n</ul>\n<p>Follow the recommended usage:\n- <strong>Login</strong> or advanced tasks in <code>on_page_context_created</code><br>\n- <strong>Custom headers</strong> or logs in <code>before_goto</code> / <code>after_goto</code><br>\n- <strong>Scrolling</strong> or final checks in <code>before_retrieve_html</code> / <code>before_return_html</code></p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/identity-based-crawling/",
    "title": "Identity Based Crawling - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"preserve-your-identity-with-crawl4ai\">Preserve Your Identity with Crawl4AI</h1>\n<p>Crawl4AI empowers you to navigate and interact with the web using your <strong>authentic digital identity</strong>, ensuring you’re recognized as a human and not mistaken for a bot. This tutorial covers:</p>\n<p>1. <strong>Managed Browsers</strong> – The recommended approach for persistent profiles and identity-based crawling.<br>\n2. <strong>Magic Mode</strong> – A simplified fallback solution for quick automation without persistent identity.</p>\n<hr>\n<h2 id=\"1-managed-browsers-your-digital-identity-solution\">1. Managed Browsers: Your Digital Identity Solution</h2>\n<p><strong>Managed Browsers</strong> let developers create and use <strong>persistent browser profiles</strong>. These profiles store local storage, cookies, and other session data, letting you browse as your <strong>real self</strong>—complete with logins, preferences, and cookies.</p>\n<h3 id=\"key-benefits\">Key Benefits</h3>\n<ul>\n<li><strong>Authentic Browsing Experience</strong>: Retain session data and browser fingerprints as though you’re a normal user.  </li>\n<li><strong>Effortless Configuration</strong>: Once you log in or solve CAPTCHAs in your chosen data directory, you can re-run crawls without repeating those steps.  </li>\n<li><strong>Empowered Data Access</strong>: If you can see the data in your own browser, you can automate its retrieval with your genuine identity.</li>\n</ul>\n<hr>\n<p>Below is a <strong>partial update</strong> to your <strong>Managed Browsers</strong> tutorial, specifically the section about <strong>creating a user-data directory</strong> using <strong>Playwright’s Chromium</strong> binary rather than a system-wide Chrome/Edge. We’ll show how to <strong>locate</strong> that binary and launch it with a <code>--user-data-dir</code> argument to set up your profile. You can then point <code>BrowserConfig.user_data_dir</code> to that folder for subsequent crawls.</p>\n<hr>\n<h3 id=\"creating-a-user-data-directory-command-line-approach-via-playwright\">Creating a User Data Directory (Command-Line Approach via Playwright)</h3>\n<p>If you installed Crawl4AI (which installs Playwright under the hood), you already have a Playwright-managed Chromium on your system. Follow these steps to launch that <strong>Chromium</strong> from your command line, specifying a <strong>custom</strong> data directory:</p>\n<p>1. <strong>Find</strong> the Playwright Chromium binary:\n   - On most systems, installed browsers go under a <code>~/.cache/ms-playwright/</code> folder or similar path.<br>\n   - To see an overview of installed browsers, run:\n     </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">python -m playwright install <span class=\"hljs-attr\">--dry-run</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n     or\n     <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">playwright install <span class=\"hljs-attr\">--dry-run</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n     (depending on your environment). This shows where Playwright keeps Chromium.<p></p>\n<ul>\n<li>For instance, you might see a path like:\n     <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">~/.cache/ms-playwright/chromium-1234/chrome-linux/chrome\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n     on Linux, or a corresponding folder on macOS/Windows.</li>\n</ul>\n<p>2. <strong>Launch</strong> the Playwright Chromium binary with a <strong>custom</strong> user-data directory:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Linux example</span>\n~/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \\\n    --user-data-dir=/home/&lt;you&gt;/my_chrome_profile\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n   <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-swift\"># macOS example (<span class=\"hljs-type\">Playwright</span>’s <span class=\"hljs-keyword\">internal</span> binary)\n<span class=\"hljs-operator\">~/</span><span class=\"hljs-type\">Library</span><span class=\"hljs-regexp\">/Caches/</span>ms<span class=\"hljs-operator\">-</span>playwright<span class=\"hljs-regexp\">/chromium-1234/</span>chrome<span class=\"hljs-operator\">-</span>mac<span class=\"hljs-regexp\">/Chromium.app/</span><span class=\"hljs-type\">Contents</span><span class=\"hljs-regexp\">/MacOS/</span><span class=\"hljs-type\">Chromium</span> \\\n    <span class=\"hljs-operator\">--</span>user<span class=\"hljs-operator\">-</span>data<span class=\"hljs-operator\">-</span>dir<span class=\"hljs-operator\">=/</span><span class=\"hljs-type\">Users</span><span class=\"hljs-regexp\">/&lt;you&gt;/</span>my_chrome_profile\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n   <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Windows example (PowerShell/cmd)</span>\n<span class=\"hljs-string\">\"C:\\Users\\&lt;you&gt;\\AppData\\Local\\ms-playwright\\chromium-1234\\chrome-win\\chrome.exe\"</span> ^\n    --user-data-<span class=\"hljs-built_in\">dir</span>=<span class=\"hljs-string\">\"C:\\Users\\&lt;you&gt;\\my_chrome_profile\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Replace</strong> the path with the actual subfolder indicated in your <code>ms-playwright</code> cache structure.<br>\n   - This <strong>opens</strong> a fresh Chromium with your new or existing data folder.<br>\n   - <strong>Log into</strong> any sites or configure your browser the way you want.<br>\n   - <strong>Close</strong> when done—your profile data is saved in that folder.</p>\n<p>3. <strong>Use</strong> that folder in <strong><code>BrowserConfig.user_data_dir</code></strong>:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\nbrowser_config = BrowserConfig(\n    headless=<span class=\"hljs-literal\">True</span>,\n    use_managed_browser=<span class=\"hljs-literal\">True</span>,\n    user_data_dir=<span class=\"hljs-string\">\"/home/&lt;you&gt;/my_chrome_profile\"</span>,\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n   - Next time you run your code, it reuses that folder—<strong>preserving</strong> your session data, cookies, local storage, etc.<p></p>\n<hr>\n<h3 id=\"creating-a-profile-using-the-crawl4ai-cli-easiest\">Creating a Profile Using the Crawl4AI CLI (Easiest)</h3>\n<p>If you prefer a guided, interactive setup, use the built-in CLI to create and manage persistent browser profiles.</p>\n<p>1.⠀Launch the profile manager:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crwl profiles\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>2.⠀Choose \"Create new profile\" and enter a profile name. A Chromium window opens so you can log in to sites and configure settings. When finished, return to the terminal and press <code>q</code> to save the profile.</p>\n<p>3.⠀Profiles are saved under <code>~/.crawl4ai/profiles/&lt;profile_name&gt;</code> (for example: <code>/home/&lt;you&gt;/.crawl4ai/profiles/test_profile_1</code>) along with a <code>storage_state.json</code> for cookies and session data.</p>\n<p>4.⠀Optionally, choose \"List profiles\" in the CLI to view available profiles and their paths.</p>\n<p>5.⠀Use the saved path with <code>BrowserConfig.user_data_dir</code>:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nprofile_path = <span class=\"hljs-string\">\"/home/&lt;you&gt;/.crawl4ai/profiles/test_profile_1\"</span>\n\nbrowser_config = BrowserConfig(\n    headless=<span class=\"hljs-literal\">True</span>,\n    use_managed_browser=<span class=\"hljs-literal\">True</span>,\n    user_data_dir=profile_path,\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com/private\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>The CLI also supports listing and deleting profiles, and even testing a crawl directly from the menu.</p>\n<hr>\n<h2 id=\"3-using-managed-browsers-in-crawl4ai\">3. Using Managed Browsers in Crawl4AI</h2>\n<p>Once you have a data directory with your session data, pass it to <strong><code>BrowserConfig</code></strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1) Reference your persistent data directory</span>\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># 'True' for automated runs</span>\n        verbose=<span class=\"hljs-literal\">True</span>,\n        use_managed_browser=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enables persistent browser strategy</span>\n        browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n        user_data_dir=<span class=\"hljs-string\">\"/path/to/my-chrome-profile\"</span>\n    )\n\n    <span class=\"hljs-comment\"># 2) Standard crawl config</span>\n    crawl_config = CrawlerRunConfig(\n        wait_for=<span class=\"hljs-string\">\"css:.logged-in-content\"</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com/private\"</span>, config=crawl_config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Successfully accessed private data with your identity!\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"workflow\">Workflow</h3>\n<p>1. <strong>Login</strong> externally (via CLI or your normal Chrome with <code>--user-data-dir=...</code>).<br>\n2. <strong>Close</strong> that browser.<br>\n3. <strong>Use</strong> the same folder in <code>user_data_dir=</code> in Crawl4AI.<br>\n4. <strong>Crawl</strong> – The site sees your identity as if you’re the same user who just logged in.</p>\n<hr>\n<h2 id=\"4-magic-mode-simplified-automation\">4. Magic Mode: Simplified Automation</h2>\n<p>If you <strong>don’t</strong> need a persistent profile or identity-based approach, <strong>Magic Mode</strong> offers a quick way to simulate human-like browsing without storing long-term data.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        config=CrawlerRunConfig(\n            magic=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Simplifies a lot of interaction</span>\n            remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n            page_timeout=<span class=\"hljs-number\">60000</span>\n        )\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Magic Mode</strong>:</p>\n<ul>\n<li>Simulates a user-like experience  </li>\n<li>Randomizes user agent &amp; navigator</li>\n<li>Randomizes interactions &amp; timings  </li>\n<li>Masks automation signals  </li>\n<li>Attempts pop-up handling  </li>\n</ul>\n<p><strong>But</strong> it’s no substitute for <strong>true</strong> user-based sessions if you want a fully legitimate identity-based solution.</p>\n<hr>\n<h2 id=\"5-comparing-managed-browsers-vs-magic-mode\">5. Comparing Managed Browsers vs. Magic Mode</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th><strong>Managed Browsers</strong></th>\n<th><strong>Magic Mode</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Session Persistence</strong></td>\n<td>Full localStorage/cookies retained in user_data_dir</td>\n<td>No persistent data (fresh each run)</td>\n</tr>\n<tr>\n<td><strong>Genuine Identity</strong></td>\n<td>Real user profile with full rights &amp; preferences</td>\n<td>Emulated user-like patterns, but no actual identity</td>\n</tr>\n<tr>\n<td><strong>Complex Sites</strong></td>\n<td>Best for login-gated sites or heavy config</td>\n<td>Simple tasks, minimal login or config needed</td>\n</tr>\n<tr>\n<td><strong>Setup</strong></td>\n<td>External creation of user_data_dir, then use in Crawl4AI</td>\n<td>Single-line approach (<code>magic=True</code>)</td>\n</tr>\n<tr>\n<td><strong>Reliability</strong></td>\n<td>Extremely consistent (same data across runs)</td>\n<td>Good for smaller tasks, can be less stable</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2 id=\"6-using-the-browserprofiler-class\">6. Using the BrowserProfiler Class</h2>\n<p>Crawl4AI provides a dedicated <code>BrowserProfiler</code> class for managing browser profiles, making it easy to create, list, and delete profiles for identity-based browsing.</p>\n<h3 id=\"creating-and-managing-profiles-with-browserprofiler\">Creating and Managing Profiles with BrowserProfiler</h3>\n<p>The <code>BrowserProfiler</code> class offers a comprehensive API for browser profile management:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> BrowserProfiler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">manage_profiles</span>():\n    <span class=\"hljs-comment\"># Create a profiler instance</span>\n    profiler = BrowserProfiler()\n\n    <span class=\"hljs-comment\"># Create a profile interactively - opens a browser window</span>\n    profile_path = <span class=\"hljs-keyword\">await</span> profiler.create_profile(\n        profile_name=<span class=\"hljs-string\">\"my-login-profile\"</span>  <span class=\"hljs-comment\"># Optional: name your profile</span>\n    )\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Profile saved at: <span class=\"hljs-subst\">{profile_path}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># List all available profiles</span>\n    profiles = profiler.list_profiles()\n\n    <span class=\"hljs-keyword\">for</span> profile <span class=\"hljs-keyword\">in</span> profiles:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Profile: <span class=\"hljs-subst\">{profile[<span class=\"hljs-string\">'name'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Path: <span class=\"hljs-subst\">{profile[<span class=\"hljs-string\">'path'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Created: <span class=\"hljs-subst\">{profile[<span class=\"hljs-string\">'created'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Browser type: <span class=\"hljs-subst\">{profile[<span class=\"hljs-string\">'type'</span>]}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Get a specific profile path by name</span>\n    specific_profile = profiler.get_profile_path(<span class=\"hljs-string\">\"my-login-profile\"</span>)\n\n    <span class=\"hljs-comment\"># Delete a profile when no longer needed</span>\n    success = profiler.delete_profile(<span class=\"hljs-string\">\"old-profile-name\"</span>)\n\nasyncio.run(manage_profiles())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>How profile creation works:</strong>\n1. A browser window opens for you to interact with\n2. You log in to websites, set preferences, etc.\n3. When you're done, press 'q' in the terminal to close the browser\n4. The profile is saved in the Crawl4AI profiles directory\n5. You can use the returned path with <code>BrowserConfig.user_data_dir</code></p>\n<h3 id=\"interactive-profile-management\">Interactive Profile Management</h3>\n<p>The <code>BrowserProfiler</code> also offers an interactive management console that guides you through profile creation, listing, and deletion:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> BrowserProfiler, AsyncWebCrawler, BrowserConfig\n\n<span class=\"hljs-comment\"># Define a function to use a profile for crawling</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_with_profile</span>(<span class=\"hljs-params\">profile_path, url</span>):\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,\n        use_managed_browser=<span class=\"hljs-literal\">True</span>,\n        user_data_dir=profile_path\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url)\n        <span class=\"hljs-keyword\">return</span> result\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Create a profiler instance</span>\n    profiler = BrowserProfiler()\n\n    <span class=\"hljs-comment\"># Launch the interactive profile manager</span>\n    <span class=\"hljs-comment\"># Passing the crawl function as a callback adds a \"crawl with profile\" option</span>\n    <span class=\"hljs-keyword\">await</span> profiler.interactive_manager(crawl_callback=crawl_with_profile)\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"legacy-methods\">Legacy Methods</h3>\n<p>For backward compatibility, the previous methods on <code>ManagedBrowser</code> are still available, but they delegate to the new <code>BrowserProfiler</code> class:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.browser_manager <span class=\"hljs-keyword\">import</span> ManagedBrowser\n\n<span class=\"hljs-comment\"># These methods still work but use BrowserProfiler internally</span>\nprofiles = ManagedBrowser.list_profiles()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"complete-example\">Complete Example</h3>\n<p>See the full example in <code>docs/examples/identity_based_browsing.py</code> for a complete demonstration of creating and using profiles for authenticated browsing using the new <code>BrowserProfiler</code> class.</p>\n<hr>\n<h2 id=\"7-locale-timezone-and-geolocation-control\">7. Locale, Timezone, and Geolocation Control</h2>\n<p>In addition to using persistent profiles, Crawl4AI supports customizing your browser's locale, timezone, and geolocation settings. These features enhance your identity-based browsing experience by allowing you to control how websites perceive your location and regional settings.</p>\n<h3 id=\"setting-locale-and-timezone\">Setting Locale and Timezone</h3>\n<p>You can set the browser's locale and timezone through <code>CrawlerRunConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai import AsyncWebCrawler, <span class=\"hljs-function\">CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        config=CrawlerRunConfig(\n            <span class=\"hljs-meta\"># Set browser locale (language and <span class=\"hljs-keyword\">region</span> formatting)</span>\n            locale=<span class=\"hljs-string\">\"fr-FR\"</span>,  <span class=\"hljs-meta\"># French (France)</span>\n\n            <span class=\"hljs-meta\"># Set browser timezone</span>\n            timezone_id=<span class=\"hljs-string\">\"Europe/Paris\"</span>,\n\n            <span class=\"hljs-meta\"># Other normal options...</span>\n            magic=True,\n            page_timeout=<span class=\"hljs-number\">60000</span>\n        )\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>How it works:</strong>\n- <code>locale</code> affects language preferences, date formats, number formats, etc.\n- <code>timezone_id</code> affects JavaScript's Date object and time-related functionality\n- These settings are applied when creating the browser context and maintained throughout the session</p>\n<h3 id=\"configuring-geolocation\">Configuring Geolocation</h3>\n<p>Control the GPS coordinates reported by the browser's geolocation API:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, GeolocationConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://maps.google.com\"</span>,  <span class=\"hljs-comment\"># Or any location-aware site</span>\n        config=CrawlerRunConfig(\n            <span class=\"hljs-comment\"># Configure precise GPS coordinates</span>\n            geolocation=GeolocationConfig(\n                latitude=<span class=\"hljs-number\">48.8566</span>,   <span class=\"hljs-comment\"># Paris coordinates</span>\n                longitude=<span class=\"hljs-number\">2.3522</span>,\n                accuracy=<span class=\"hljs-number\">100</span>        <span class=\"hljs-comment\"># Accuracy in meters (optional)</span>\n            ),\n\n            <span class=\"hljs-comment\"># This site will see you as being in Paris</span>\n            page_timeout=<span class=\"hljs-number\">60000</span>\n        )\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Important notes:</strong>\n- When <code>geolocation</code> is specified, the browser is automatically granted permission to access location\n- Websites using the Geolocation API will receive the exact coordinates you specify\n- This affects map services, store locators, delivery services, etc.\n- Combined with the appropriate <code>locale</code> and <code>timezone_id</code>, you can create a fully consistent location profile</p>\n<h3 id=\"combining-with-managed-browsers\">Combining with Managed Browsers</h3>\n<p>These settings work perfectly with managed browsers for a complete identity solution:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-title\">import</span> (<span class=\"hljs-params\">\n    AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, \n    GeolocationConfig\n</span>)\n\nbrowser_config</span> = BrowserConfig(\n    use_managed_browser=True,\n    user_data_dir=<span class=\"hljs-string\">\"/path/to/my-profile\"</span>,\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>\n)\n\ncrawl_config = CrawlerRunConfig(\n    <span class=\"hljs-meta\"># Location settings</span>\n    locale=<span class=\"hljs-string\">\"es-MX\"</span>,                  <span class=\"hljs-meta\"># Spanish (Mexico)</span>\n    timezone_id=<span class=\"hljs-string\">\"America/Mexico_City\"</span>,\n    geolocation=GeolocationConfig(\n        latitude=<span class=\"hljs-number\">19.4326</span>,            <span class=\"hljs-meta\"># Mexico City</span>\n        longitude=<span class=\"hljs-number\">-99.1332</span>\n    )\n)\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>(<span class=\"hljs-params\">config=browser_config</span>) <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=crawl_config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Combining persistent profiles with precise geolocation and region settings gives you complete control over your digital identity.</p>\n<h2 id=\"8-summary\">8. Summary</h2>\n<ul>\n<li><strong>Create</strong> your user-data directory either:</li>\n<li>By launching Chrome/Chromium externally with <code>--user-data-dir=/some/path</code> </li>\n<li>Or by using the built-in <code>BrowserProfiler.create_profile()</code> method</li>\n<li>Or through the interactive interface with <code>profiler.interactive_manager()</code></li>\n<li><strong>Log in</strong> or configure sites as needed, then close the browser</li>\n<li><strong>Reference</strong> that folder in <code>BrowserConfig(user_data_dir=\"...\")</code> + <code>use_managed_browser=True</code></li>\n<li><strong>Customize</strong> identity aspects with <code>locale</code>, <code>timezone_id</code>, and <code>geolocation</code></li>\n<li><strong>List and reuse</strong> profiles with <code>BrowserProfiler.list_profiles()</code></li>\n<li><strong>Manage</strong> your profiles with the dedicated <code>BrowserProfiler</code> class</li>\n<li>Enjoy <strong>persistent</strong> sessions that reflect your real identity</li>\n<li>If you only need quick, ephemeral automation, <strong>Magic Mode</strong> might suffice</li>\n</ul>\n<p><strong>Recommended</strong>: Always prefer a <strong>Managed Browser</strong> for robust, identity-based crawling and simpler interactions with complex sites. Use <strong>Magic Mode</strong> for quick tasks or prototypes where persistent data is unnecessary.</p>\n<p>With these approaches, you preserve your <strong>authentic</strong> browsing environment, ensuring the site sees you exactly as a normal user—no repeated logins or wasted time.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/lazy-loading/",
    "title": "Lazy Loading - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h2 id=\"handling-lazy-loaded-images\">Handling Lazy-Loaded Images</h2>\n<p>Many websites now load images <strong>lazily</strong> as you scroll. If you need to ensure they appear in your final crawl (and in <code>result.media</code>), consider:</p>\n<p>1. <strong><code>wait_for_images=True</code></strong> – Wait for images to fully load.<br>\n2. <strong><code>scan_full_page</code></strong> – Force the crawler to scroll the entire page, triggering lazy loads.<br>\n3. <strong><code>scroll_delay</code></strong> – Add small delays between scroll steps.  </p>\n<p><strong>Note</strong>: If the site requires multiple “Load More” triggers or complex interactions, see the <a href=\"../../core/page-interaction/\">Page Interaction docs</a>. For sites with virtual scrolling (Twitter/Instagram style), see the <a href=\"../virtual-scroll/\">Virtual Scroll docs</a>.</p>\n<h3 id=\"example-ensuring-lazy-images-appear\">Example: Ensuring Lazy Images Appear</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, BrowserConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Force the crawler to wait until images are fully loaded</span>\n        wait_for_images=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Option 1: If you want to automatically scroll the page to load images</span>\n        scan_full_page=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Tells the crawler to try scrolling the entire page</span>\n        scroll_delay=<span class=\"hljs-number\">0.5</span>,     <span class=\"hljs-comment\"># Delay (seconds) between scroll steps</span>\n\n        <span class=\"hljs-comment\"># Option 2: If the site uses a 'Load More' or JS triggers for images,</span>\n        <span class=\"hljs-comment\"># you can also specify js_code or wait_for logic here.</span>\n\n        cache_mode=CacheMode.BYPASS,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://www.example.com/gallery\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            images = result.media.get(<span class=\"hljs-string\">\"images\"</span>, [])\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Images found:\"</span>, <span class=\"hljs-built_in\">len</span>(images))\n            <span class=\"hljs-keyword\">for</span> i, img <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(images[:<span class=\"hljs-number\">5</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[Image <span class=\"hljs-subst\">{i}</span>] URL: <span class=\"hljs-subst\">{img[<span class=\"hljs-string\">'src'</span>]}</span>, Score: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'score'</span>,<span class=\"hljs-string\">'N/A'</span>)}</span>\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Explanation</strong>:</p>\n<ul>\n<li><strong><code>wait_for_images=True</code></strong><br>\n  The crawler tries to ensure images have finished loading before finalizing the HTML.  </li>\n<li><strong><code>scan_full_page=True</code></strong><br>\n  Tells the crawler to attempt scrolling from top to bottom. Each scroll step helps trigger lazy loading.  </li>\n<li><strong><code>scroll_delay=0.5</code></strong><br>\n  Pause half a second between each scroll step. Helps the site load images before continuing.</li>\n</ul>\n<p><strong>When to Use</strong>:</p>\n<ul>\n<li><strong>Lazy-Loading</strong>: If images appear only when the user scrolls into view, <code>scan_full_page</code> + <code>scroll_delay</code> helps the crawler see them.  </li>\n<li><strong>Heavier Pages</strong>: If a page is extremely long, be mindful that scanning the entire page can be slow. Adjust <code>scroll_delay</code> or the max scroll steps as needed.</li>\n</ul>\n<hr>\n<h2 id=\"combining-with-other-link-media-filters\">Combining with Other Link &amp; Media Filters</h2>\n<p>You can still combine <strong>lazy-load</strong> logic with the usual <strong>exclude_external_images</strong>, <strong>exclude_domains</strong>, or link filtration:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    wait_for_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    scan_full_page<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    scroll_delay<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.5</span>,\n\n    <span class=\"hljs-comment\"># Filter out external images if you only want local ones</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n\n    <span class=\"hljs-comment\"># Exclude certain domains for links</span>\n    exclude_domains<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"spammycdn.com\"</span><span class=\"hljs-punctuation\">]</span>,\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This approach ensures you see <strong>all</strong> images from the main domain while ignoring external ones, and the crawler physically scrolls the entire page so that lazy-loading triggers.</p>\n<hr>\n<h2 id=\"tips-troubleshooting\">Tips &amp; Troubleshooting</h2>\n<p>1. <strong>Long Pages</strong><br>\n   - Setting <code>scan_full_page=True</code> on extremely long or infinite-scroll pages can be resource-intensive.<br>\n   - Consider using <a href=\"../../core/page-interaction/\">hooks</a> or specialized logic to load specific sections or “Load More” triggers repeatedly.</p>\n<p>2. <strong>Mixed Image Behavior</strong><br>\n   - Some sites load images in batches as you scroll. If you’re missing images, increase your <code>scroll_delay</code> or call multiple partial scrolls in a loop with JS code or hooks.</p>\n<p>3. <strong>Combining with Dynamic Wait</strong><br>\n   - If the site has a placeholder that only changes to a real image after a certain event, you might do <code>wait_for=\"css:img.loaded\"</code> or a custom JS <code>wait_for</code>.</p>\n<p>4. <strong>Caching</strong><br>\n   - If <code>cache_mode</code> is enabled, repeated crawls might skip some network fetches. If you suspect caching is missing new images, set <code>cache_mode=CacheMode.BYPASS</code> for fresh fetches.</p>\n<hr>\n<p>With <strong>lazy-loading</strong> support, <strong>wait_for_images</strong>, and <strong>scan_full_page</strong> settings, you can capture the entire gallery or feed of images you expect—even if the site only loads them as the user scrolls. Combine these with the standard media filtering and domain exclusion for a complete link &amp; media handling strategy.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/multi-url-crawling/",
    "title": "Multi-URL Crawling - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"advanced-multi-url-crawling-with-dispatchers\">Advanced Multi-URL Crawling with Dispatchers</h1>\n<blockquote>\n<p><strong>Heads Up</strong>: Crawl4AI supports advanced dispatchers for <strong>parallel</strong> or <strong>throttled</strong> crawling, providing dynamic rate limiting and memory usage checks. The built-in <code>arun_many()</code> function uses these dispatchers to handle concurrency efficiently.</p>\n</blockquote>\n<h2 id=\"1-introduction\">1. Introduction</h2>\n<p>When crawling many URLs:</p>\n<ul>\n<li><strong>Basic</strong>: Use <code>arun()</code> in a loop (simple but less efficient)</li>\n<li><strong>Better</strong>: Use <code>arun_many()</code>, which efficiently handles multiple URLs with proper concurrency control</li>\n<li><strong>Best</strong>: Customize dispatcher behavior for your specific needs (memory management, rate limits, etc.)</li>\n</ul>\n<p><strong>Why Dispatchers?</strong>  </p>\n<ul>\n<li><strong>Adaptive</strong>: Memory-based dispatchers can pause or slow down based on system resources</li>\n<li><strong>Rate-limiting</strong>: Built-in rate limiting with exponential backoff for 429/503 responses</li>\n<li><strong>Real-time Monitoring</strong>: Live dashboard of ongoing tasks, memory usage, and performance</li>\n<li><strong>Flexibility</strong>: Choose between memory-adaptive or semaphore-based concurrency</li>\n</ul>\n<hr>\n<h2 id=\"2-core-components\">2. Core Components</h2>\n<h3 id=\"21-rate-limiter\">2.1 Rate Limiter</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">RateLimiter</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        <span class=\"hljs-comment\"># Random delay range between requests</span>\n        base_delay: <span class=\"hljs-type\">Tuple</span>[<span class=\"hljs-built_in\">float</span>, <span class=\"hljs-built_in\">float</span>] = (<span class=\"hljs-params\"><span class=\"hljs-number\">1.0</span>, <span class=\"hljs-number\">3.0</span></span>),  \n\n        <span class=\"hljs-comment\"># Maximum backoff delay</span>\n        max_delay: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">60.0</span>,                        \n\n        <span class=\"hljs-comment\"># Retries before giving up</span>\n        max_retries: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">3</span>,                          \n\n        <span class=\"hljs-comment\"># Status codes triggering backoff</span>\n        rate_limit_codes: <span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">int</span>] = [<span class=\"hljs-number\">429</span>, <span class=\"hljs-number\">503</span>]        \n    </span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Here’s the revised and simplified explanation of the <strong>RateLimiter</strong>, focusing on constructor parameters and adhering to your markdown style and mkDocs guidelines.</p>\n<h4 id=\"ratelimiter-constructor-parameters\">RateLimiter Constructor Parameters</h4>\n<p>The <strong>RateLimiter</strong> is a utility that helps manage the pace of requests to avoid overloading servers or getting blocked due to rate limits. It operates internally to delay requests and handle retries but can be configured using its constructor parameters.</p>\n<p><strong>Parameters of the <code>RateLimiter</code> constructor:</strong></p>\n<p>1. <strong><code>base_delay</code></strong> (<code>Tuple[float, float]</code>, default: <code>(1.0, 3.0)</code>)<br>\n  The range for a random delay (in seconds) between consecutive requests to the same domain.</p>\n<ul>\n<li>A random delay is chosen between <code>base_delay[0]</code> and <code>base_delay[1]</code> for each request.  </li>\n<li>This prevents sending requests at a predictable frequency, reducing the chances of triggering rate limits.</li>\n</ul>\n<p><strong>Example:</strong><br>\nIf <code>base_delay = (2.0, 5.0)</code>, delays could be randomly chosen as <code>2.3s</code>, <code>4.1s</code>, etc.</p>\n<hr>\n<p>2. <strong><code>max_delay</code></strong> (<code>float</code>, default: <code>60.0</code>)<br>\n  The maximum allowable delay when rate-limiting errors occur.</p>\n<ul>\n<li>When servers return rate-limit responses (e.g., 429 or 503), the delay increases exponentially with jitter.  </li>\n<li>The <code>max_delay</code> ensures the delay doesn’t grow unreasonably high, capping it at this value.</li>\n</ul>\n<p><strong>Example:</strong><br>\nFor a <code>max_delay = 30.0</code>, even if backoff calculations suggest a delay of <code>45s</code>, it will cap at <code>30s</code>.</p>\n<hr>\n<p>3. <strong><code>max_retries</code></strong> (<code>int</code>, default: <code>3</code>)<br>\n  The maximum number of retries for a request if rate-limiting errors occur.</p>\n<ul>\n<li>After encountering a rate-limit response, the <code>RateLimiter</code> retries the request up to this number of times.  </li>\n<li>If all retries fail, the request is marked as failed, and the process continues.</li>\n</ul>\n<p><strong>Example:</strong><br>\nIf <code>max_retries = 3</code>, the system retries a failed request three times before giving up.</p>\n<hr>\n<p>4. <strong><code>rate_limit_codes</code></strong> (<code>List[int]</code>, default: <code>[429, 503]</code>)<br>\n  A list of HTTP status codes that trigger the rate-limiting logic.</p>\n<ul>\n<li>These status codes indicate the server is overwhelmed or actively limiting requests.  </li>\n<li>You can customize this list to include other codes based on specific server behavior.</li>\n</ul>\n<p><strong>Example:</strong><br>\nIf <code>rate_limit_codes = [429, 503, 504]</code>, the crawler will back off on these three error codes.</p>\n<hr>\n<p><strong>How to Use the <code>RateLimiter</code>:</strong></p>\n<p>Here’s an example of initializing and using a <code>RateLimiter</code> in your project:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">from crawl4ai import RateLimiter\n\n<span class=\"hljs-comment\"># Create a RateLimiter with custom settings</span>\nrate_limiter <span class=\"hljs-punctuation\">=</span> RateLimiter<span class=\"hljs-punctuation\">(</span>\n    base_delay<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">(</span><span class=\"hljs-number\">2.0</span>, <span class=\"hljs-number\">4.0</span><span class=\"hljs-punctuation\">)</span>,  <span class=\"hljs-comment\"># Random delay between 2-4 seconds</span>\n    max_delay<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">30.0</span>,         <span class=\"hljs-comment\"># Cap delay at 30 seconds</span>\n    max_retries<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5</span>,          <span class=\"hljs-comment\"># Retry up to 5 times on rate-limiting errors</span>\n    rate_limit_codes<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-number\">429</span>, <span class=\"hljs-number\">503</span><span class=\"hljs-punctuation\">]</span>  <span class=\"hljs-comment\"># Handle these HTTP status codes</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># RateLimiter will handle delays and retries internally</span>\n<span class=\"hljs-comment\"># No additional setup is required for its operation</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The <code>RateLimiter</code> integrates seamlessly with dispatchers like <code>MemoryAdaptiveDispatcher</code> and <code>SemaphoreDispatcher</code>, ensuring requests are paced correctly without user intervention. Its internal mechanisms manage delays and retries to avoid overwhelming servers while maximizing efficiency.</p>\n<h3 id=\"22-crawler-monitor\">2.2 Crawler Monitor</h3>\n<p>The CrawlerMonitor provides real-time visibility into crawling operations:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerMonitor, DisplayMode\nmonitor = CrawlerMonitor(\n    <span class=\"hljs-comment\"># Maximum rows in live display</span>\n    max_visible_rows=<span class=\"hljs-number\">15</span>,          \n\n    <span class=\"hljs-comment\"># DETAILED or AGGREGATED view</span>\n    display_mode=DisplayMode.DETAILED  \n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Display Modes</strong>:</p>\n<ol>\n<li><strong>DETAILED</strong>: Shows individual task status, memory usage, and timing</li>\n<li><strong>AGGREGATED</strong>: Displays summary statistics and overall progress</li>\n</ol>\n<hr>\n<h2 id=\"3-available-dispatchers\">3. Available Dispatchers</h2>\n<h3 id=\"31-memoryadaptivedispatcher-default\">3.1 MemoryAdaptiveDispatcher (Default)</h3>\n<p>Automatically manages concurrency based on system memory usage:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_dispatcher <span class=\"hljs-keyword\">import</span> MemoryAdaptiveDispatcher\n\ndispatcher = MemoryAdaptiveDispatcher(\n    memory_threshold_percent=<span class=\"hljs-number\">90.0</span>,  <span class=\"hljs-comment\"># Pause if memory exceeds this</span>\n    check_interval=<span class=\"hljs-number\">1.0</span>,             <span class=\"hljs-comment\"># How often to check memory</span>\n    max_session_permit=<span class=\"hljs-number\">10</span>,          <span class=\"hljs-comment\"># Maximum concurrent tasks</span>\n    rate_limiter=RateLimiter(       <span class=\"hljs-comment\"># Optional rate limiting</span>\n        base_delay=(<span class=\"hljs-number\">1.0</span>, <span class=\"hljs-number\">2.0</span>),\n        max_delay=<span class=\"hljs-number\">30.0</span>,\n        max_retries=<span class=\"hljs-number\">2</span>\n    ),\n    monitor=CrawlerMonitor(         <span class=\"hljs-comment\"># Optional monitoring</span>\n        max_visible_rows=<span class=\"hljs-number\">15</span>,\n        display_mode=DisplayMode.DETAILED\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Constructor Parameters:</strong></p>\n<p>1. <strong><code>memory_threshold_percent</code></strong> (<code>float</code>, default: <code>90.0</code>)<br>\n  Specifies the memory usage threshold (as a percentage). If system memory usage exceeds this value, the dispatcher pauses crawling to prevent system overload.</p>\n<p>2. <strong><code>check_interval</code></strong> (<code>float</code>, default: <code>1.0</code>)<br>\n  The interval (in seconds) at which the dispatcher checks system memory usage.</p>\n<p>3. <strong><code>max_session_permit</code></strong> (<code>int</code>, default: <code>10</code>)<br>\n  The maximum number of concurrent crawling tasks allowed. This ensures resource limits are respected while maintaining concurrency.</p>\n<p>4. <strong><code>memory_wait_timeout</code></strong> (<code>float</code>, default: <code>600.0</code>)\n  Optional timeout (in seconds). If memory usage exceeds <code>memory_threshold_percent</code> for longer than this duration, a <code>MemoryError</code> is raised.</p>\n<p>5. <strong><code>rate_limiter</code></strong> (<code>RateLimiter</code>, default: <code>None</code>)<br>\n  Optional rate-limiting logic to avoid server-side blocking (e.g., for handling 429 or 503 errors). See <strong>RateLimiter</strong> for details.</p>\n<p>6. <strong><code>monitor</code></strong> (<code>CrawlerMonitor</code>, default: <code>None</code>)<br>\n  Optional monitoring for real-time task tracking and performance insights. See <strong>CrawlerMonitor</strong> for details.</p>\n<hr>\n<h3 id=\"32-semaphoredispatcher\">3.2 SemaphoreDispatcher</h3>\n<p>Provides simple concurrency control with a fixed limit:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_dispatcher <span class=\"hljs-keyword\">import</span> SemaphoreDispatcher\n\ndispatcher = SemaphoreDispatcher(\n    max_session_permit=<span class=\"hljs-number\">20</span>,         <span class=\"hljs-comment\"># Maximum concurrent tasks</span>\n    rate_limiter=RateLimiter(      <span class=\"hljs-comment\"># Optional rate limiting</span>\n        base_delay=(<span class=\"hljs-number\">0.5</span>, <span class=\"hljs-number\">1.0</span>),\n        max_delay=<span class=\"hljs-number\">10.0</span>\n    ),\n    monitor=CrawlerMonitor(        <span class=\"hljs-comment\"># Optional monitoring</span>\n        max_visible_rows=<span class=\"hljs-number\">15</span>,\n        display_mode=DisplayMode.DETAILED\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Constructor Parameters:</strong></p>\n<p>1. <strong><code>max_session_permit</code></strong> (<code>int</code>, default: <code>20</code>)<br>\n  The maximum number of concurrent crawling tasks allowed, irrespective of semaphore slots.</p>\n<p>2. <strong><code>rate_limiter</code></strong> (<code>RateLimiter</code>, default: <code>None</code>)<br>\n  Optional rate-limiting logic to avoid overwhelming servers. See <strong>RateLimiter</strong> for details.</p>\n<p>3. <strong><code>monitor</code></strong> (<code>CrawlerMonitor</code>, default: <code>None</code>)<br>\n  Optional monitoring for tracking task progress and resource usage. See <strong>CrawlerMonitor</strong> for details.</p>\n<hr>\n<h2 id=\"4-usage-examples\">4. Usage Examples</h2>\n<h3 id=\"41-batch-processing-default\">4.1 Batch Processing (Default)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_batch</span>():\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, verbose=<span class=\"hljs-literal\">False</span>)\n    run_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        stream=<span class=\"hljs-literal\">False</span>  <span class=\"hljs-comment\"># Default: get all results at once</span>\n    )\n\n    dispatcher = MemoryAdaptiveDispatcher(\n        memory_threshold_percent=<span class=\"hljs-number\">70.0</span>,\n        check_interval=<span class=\"hljs-number\">1.0</span>,\n        max_session_permit=<span class=\"hljs-number\">10</span>,\n        monitor=CrawlerMonitor(\n            display_mode=DisplayMode.DETAILED\n        )\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Get all results at once</span>\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            urls=urls,\n            config=run_config,\n            dispatcher=dispatcher\n        )\n\n        <span class=\"hljs-comment\"># Process all results after completion</span>\n        <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-keyword\">await</span> process_result(result)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl <span class=\"hljs-subst\">{result.url}</span>: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Review:</strong><br>\n- <strong>Purpose:</strong> Executes a batch crawl with all URLs processed together after crawling is complete.<br>\n- <strong>Dispatcher:</strong> Uses <code>MemoryAdaptiveDispatcher</code> to manage concurrency and system memory.<br>\n- <strong>Stream:</strong> Disabled (<code>stream=False</code>), so all results are collected at once for post-processing.<br>\n- <strong>Best Use Case:</strong> When you need to analyze results in bulk rather than individually during the crawl.</p>\n<hr>\n<h3 id=\"42-streaming-mode\">4.2 Streaming Mode</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_streaming</span>():\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, verbose=<span class=\"hljs-literal\">False</span>)\n    run_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        stream=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    )\n\n    dispatcher = MemoryAdaptiveDispatcher(\n        memory_threshold_percent=<span class=\"hljs-number\">70.0</span>,\n        check_interval=<span class=\"hljs-number\">1.0</span>,\n        max_session_permit=<span class=\"hljs-number\">10</span>,\n        monitor=CrawlerMonitor(\n            display_mode=DisplayMode.DETAILED\n        )\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Process results as they become available</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            urls=urls,\n            config=run_config,\n            dispatcher=dispatcher\n        ):\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-comment\"># Process each result immediately</span>\n                <span class=\"hljs-keyword\">await</span> process_result(result)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl <span class=\"hljs-subst\">{result.url}</span>: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Review:</strong><br>\n- <strong>Purpose:</strong> Enables streaming to process results as soon as they’re available.<br>\n- <strong>Dispatcher:</strong> Uses <code>MemoryAdaptiveDispatcher</code> for concurrency and memory management.<br>\n- <strong>Stream:</strong> Enabled (<code>stream=True</code>), allowing real-time processing during crawling.<br>\n- <strong>Best Use Case:</strong> When you need to act on results immediately, such as for real-time analytics or progressive data storage.</p>\n<hr>\n<h3 id=\"43-semaphore-based-crawling\">4.3 Semaphore-based Crawling</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_with_semaphore</span>(<span class=\"hljs-params\">urls</span>):\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, verbose=<span class=\"hljs-literal\">False</span>)\n    run_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n\n    dispatcher = SemaphoreDispatcher(\n        semaphore_count=<span class=\"hljs-number\">5</span>,\n        rate_limiter=RateLimiter(\n            base_delay=(<span class=\"hljs-number\">0.5</span>, <span class=\"hljs-number\">1.0</span>),\n            max_delay=<span class=\"hljs-number\">10.0</span>\n        ),\n        monitor=CrawlerMonitor(\n            max_visible_rows=<span class=\"hljs-number\">15</span>,\n            display_mode=DisplayMode.DETAILED\n        )\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            urls, \n            config=run_config,\n            dispatcher=dispatcher\n        )\n        <span class=\"hljs-keyword\">return</span> results\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Review:</strong><br>\n- <strong>Purpose:</strong> Uses <code>SemaphoreDispatcher</code> to limit concurrency with a fixed number of slots.<br>\n- <strong>Dispatcher:</strong> Configured with a semaphore to control parallel crawling tasks.<br>\n- <strong>Rate Limiter:</strong> Prevents servers from being overwhelmed by pacing requests.<br>\n- <strong>Best Use Case:</strong> When you want precise control over the number of concurrent requests, independent of system memory.</p>\n<hr>\n<h3 id=\"44-robotstxt-consideration\">4.4 Robots.txt Consideration</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    urls = [\n        <span class=\"hljs-string\">\"https://example1.com\"</span>,\n        <span class=\"hljs-string\">\"https://example2.com\"</span>,\n        <span class=\"hljs-string\">\"https://example3.com\"</span>\n    ]\n\n    config = CrawlerRunConfig(\n        cache_mode=CacheMode.ENABLED,\n        check_robots_txt=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Will respect robots.txt for each URL</span>\n        semaphore_count=<span class=\"hljs-number\">3</span>      <span class=\"hljs-comment\"># Max concurrent requests</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> crawler.arun_many(urls, config=config):\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully crawled <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n            <span class=\"hljs-keyword\">elif</span> result.status_code == <span class=\"hljs-number\">403</span> <span class=\"hljs-keyword\">and</span> <span class=\"hljs-string\">\"robots.txt\"</span> <span class=\"hljs-keyword\">in</span> result.error_message:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Skipped <span class=\"hljs-subst\">{result.url}</span> - blocked by robots.txt\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl <span class=\"hljs-subst\">{result.url}</span>: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Review:</strong><br>\n- <strong>Purpose:</strong> Ensures compliance with <code>robots.txt</code> rules for ethical and legal web crawling.<br>\n- <strong>Configuration:</strong> Set <code>check_robots_txt=True</code> to validate each URL against <code>robots.txt</code> before crawling.<br>\n- <strong>Dispatcher:</strong> Handles requests with concurrency limits (<code>semaphore_count=3</code>).<br>\n- <strong>Best Use Case:</strong> When crawling websites that strictly enforce robots.txt policies or for responsible crawling practices.</p>\n<hr>\n<h2 id=\"5-dispatch-results\">5. Dispatch Results</h2>\n<p>Each crawl result includes dispatch information:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\"><span class=\"hljs-keyword\">@dataclass</span>\nclass <span class=\"hljs-attribute\">DispatchResult</span>:\n    task_<span class=\"hljs-attribute\">id</span>: str\n    memory_<span class=\"hljs-attribute\">usage</span>: float\n    peak_<span class=\"hljs-attribute\">memory</span>: float\n    start_<span class=\"hljs-attribute\">time</span>: datetime\n    end_<span class=\"hljs-attribute\">time</span>: datetime\n    error_<span class=\"hljs-attribute\">message</span>: str = <span class=\"hljs-string\">\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Access via <code>result.dispatch_result</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n    <span class=\"hljs-keyword\">if</span> result.success:\n        dr = result.dispatch_result\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"URL: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Memory: <span class=\"hljs-subst\">{dr.memory_usage:<span class=\"hljs-number\">.1</span>f}</span>MB\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Duration: <span class=\"hljs-subst\">{dr.end_time - dr.start_time}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"6-url-specific-configurations\">6. URL-Specific Configurations</h2>\n<p>When crawling diverse content types, you often need different configurations for different URLs. For example:\n- PDFs need specialized extraction\n- Blog pages benefit from content filtering\n- Dynamic sites need JavaScript execution\n- API endpoints need JSON parsing</p>\n<h3 id=\"61-basic-url-pattern-matching\">6.1 Basic URL Pattern Matching</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, MatchMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.extraction_strategy <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_mixed_content</span>():\n    <span class=\"hljs-comment\"># Configure different strategies for different content</span>\n    configs = [\n        <span class=\"hljs-comment\"># PDF files - specialized extraction</span>\n        CrawlerRunConfig(\n            url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>,\n            scraping_strategy=PDFContentScrapingStrategy()\n        ),\n\n        <span class=\"hljs-comment\"># Blog/article pages - content filtering</span>\n        CrawlerRunConfig(\n            url_matcher=[<span class=\"hljs-string\">\"*/blog/*\"</span>, <span class=\"hljs-string\">\"*/article/*\"</span>],\n            markdown_generator=DefaultMarkdownGenerator(\n                content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.48</span>)\n            )\n        ),\n\n        <span class=\"hljs-comment\"># Dynamic pages - JavaScript execution</span>\n        CrawlerRunConfig(\n            url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'github.com'</span> <span class=\"hljs-keyword\">in</span> url,\n            js_code=<span class=\"hljs-string\">\"window.scrollTo(0, 500);\"</span>\n        ),\n\n        <span class=\"hljs-comment\"># API endpoints - JSON extraction</span>\n        CrawlerRunConfig(\n            url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'api'</span> <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">or</span> url.endswith(<span class=\"hljs-string\">'.json'</span>),\n            <span class=\"hljs-comment\"># Custome settings for JSON extraction</span>\n        ),\n\n        <span class=\"hljs-comment\"># Default config for everything else</span>\n        CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher means it matches ALL URLs (fallback)</span>\n    ]\n\n    <span class=\"hljs-comment\"># Mixed URLs</span>\n    urls = [\n        <span class=\"hljs-string\">\"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf\"</span>,\n        <span class=\"hljs-string\">\"https://blog.python.org/\"</span>,\n        <span class=\"hljs-string\">\"https://github.com/microsoft/playwright\"</span>,\n        <span class=\"hljs-string\">\"https://httpbin.org/json\"</span>,\n        <span class=\"hljs-string\">\"https://example.com/\"</span>\n    ]\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            urls=urls,\n            config=configs  <span class=\"hljs-comment\"># Pass list of configs</span>\n        )\n\n        <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{result.url}</span>: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.markdown)}</span> chars\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"62-advanced-pattern-matching\">6.2 Advanced Pattern Matching</h3>\n<p><strong>Important</strong>: A <code>CrawlerRunConfig</code> without <code>url_matcher</code> (or with <code>url_matcher=None</code>) matches ALL URLs. This makes it perfect as a default/fallback configuration.</p>\n<p>The <code>url_matcher</code> parameter supports three types of patterns:</p>\n<h4 id=\"glob-patterns-strings\">Glob Patterns (Strings)</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Simple patterns</span>\n<span class=\"hljs-string\">\"*.pdf\"</span>                    <span class=\"hljs-comment\"># Any PDF file</span>\n<span class=\"hljs-string\">\"*/api/*\"</span>                  <span class=\"hljs-comment\"># Any URL with /api/ in path</span>\n<span class=\"hljs-string\">\"https://*.example.com/*\"</span>  <span class=\"hljs-comment\"># Subdomain matching</span>\n<span class=\"hljs-string\">\"*://example.com/blog/*\"</span>   <span class=\"hljs-comment\"># Any protocol</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"custom-functions\">Custom Functions</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Complex logic with lambdas</span>\n<span class=\"hljs-keyword\">lambda</span> url: url.startswith(<span class=\"hljs-string\">'https://'</span>) <span class=\"hljs-keyword\">and</span> <span class=\"hljs-string\">'secure'</span> <span class=\"hljs-keyword\">in</span> url\n<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-built_in\">len</span>(url) &gt; <span class=\"hljs-number\">50</span> <span class=\"hljs-keyword\">and</span> url.count(<span class=\"hljs-string\">'/'</span>) &gt; <span class=\"hljs-number\">5</span>\n<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-built_in\">any</span>(domain <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">for</span> domain <span class=\"hljs-keyword\">in</span> [<span class=\"hljs-string\">'api.'</span>, <span class=\"hljs-string\">'data.'</span>, <span class=\"hljs-string\">'feed.'</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"mixed-lists-with-andor-logic\">Mixed Lists with AND/OR Logic</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Combine multiple conditions</span>\nCrawlerRunConfig(\n    url_matcher=[\n        <span class=\"hljs-string\">\"https://*\"</span>,                        <span class=\"hljs-comment\"># Must be HTTPS</span>\n        <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'internal'</span> <span class=\"hljs-keyword\">in</span> url,      <span class=\"hljs-comment\"># Must contain 'internal'</span>\n        <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-keyword\">not</span> url.endswith(<span class=\"hljs-string\">'.pdf'</span>) <span class=\"hljs-comment\"># Must not be PDF</span>\n    ],\n    match_mode=MatchMode.AND  <span class=\"hljs-comment\"># ALL conditions must match</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"63-practical-example-news-site-crawler\">6.3 Practical Example: News Site Crawler</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> def <span class=\"hljs-title\">crawl_news_site</span>():\n    dispatcher</span> = MemoryAdaptiveDispatcher(\n        memory_threshold_percent=<span class=\"hljs-number\">70.0</span>,\n        rate_limiter=RateLimiter(base_delay=(<span class=\"hljs-number\">1.0</span>, <span class=\"hljs-number\">2.0</span>))\n    )\n\n    configs = [\n        <span class=\"hljs-meta\"># Homepage - light extraction</span>\n        CrawlerRunConfig(\n            url_matcher=lambda url: url.rstrip(<span class=\"hljs-string\">'/'</span>) == <span class=\"hljs-string\">'https://news.ycombinator.com'</span>,\n            css_selector=<span class=\"hljs-string\">\"nav, .headline\"</span>,\n            extraction_strategy=None\n        ),\n\n        <span class=\"hljs-meta\"># Article pages - full extraction</span>\n        CrawlerRunConfig(\n            url_matcher=<span class=\"hljs-string\">\"*/article/*\"</span>,\n            extraction_strategy=CosineStrategy(\n                semantic_filter=<span class=\"hljs-string\">\"article content\"</span>,\n                word_count_threshold=<span class=\"hljs-number\">100</span>\n            ),\n            screenshot=True,\n            excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"aside\"</span>, <span class=\"hljs-string\">\"footer\"</span>]\n        ),\n\n        <span class=\"hljs-meta\"># Author pages - metadata focus</span>\n        CrawlerRunConfig(\n            url_matcher=<span class=\"hljs-string\">\"*/author/*\"</span>,\n            extraction_strategy=JsonCssExtractionStrategy({\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"h1.author-name\"</span>,\n                <span class=\"hljs-string\">\"bio\"</span>: <span class=\"hljs-string\">\".author-bio\"</span>,\n                <span class=\"hljs-string\">\"articles\"</span>: <span class=\"hljs-string\">\"article.post-card h2\"</span>\n            })\n        ),\n\n        <span class=\"hljs-meta\"># Everything <span class=\"hljs-keyword\">else</span></span>\n        CrawlerRunConfig()\n    ]\n\n    <span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n        results</span> = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            urls=news_urls,\n            config=configs,\n            dispatcher=dispatcher\n        )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"64-best-practices\">6.4 Best Practices</h3>\n<ol>\n<li><strong>Order Matters</strong>: Configs are evaluated in order - put specific patterns before general ones</li>\n<li><strong>Default Config Behavior</strong>: </li>\n<li>A config without <code>url_matcher</code> matches ALL URLs</li>\n<li>Always include a default config as the last item if you want to handle all URLs</li>\n<li>Without a default config, unmatched URLs will fail with \"No matching configuration found\"</li>\n<li><strong>Test Your Patterns</strong>: Use the config's <code>is_match()</code> method to test patterns:\n   <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">config = CrawlerRunConfig(url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>)\n<span class=\"hljs-built_in\">print</span>(config.is_match(<span class=\"hljs-string\">\"https://example.com/doc.pdf\"</span>))  <span class=\"hljs-comment\"># True</span>\n\ndefault_config = CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher</span>\n<span class=\"hljs-built_in\">print</span>(default_config.is_match(<span class=\"hljs-string\">\"https://any-url.com\"</span>))  <span class=\"hljs-comment\"># True - matches everything!</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li><strong>Optimize for Performance</strong>: </li>\n<li>Disable JS for static content</li>\n<li>Skip screenshots for data APIs</li>\n<li>Use appropriate extraction strategies</li>\n</ol>\n<h2 id=\"7-summary\">7. Summary</h2>\n<p>1. <strong>Two Dispatcher Types</strong>:</p>\n<ul>\n<li>MemoryAdaptiveDispatcher (default): Dynamic concurrency based on memory</li>\n<li>SemaphoreDispatcher: Fixed concurrency limit</li>\n</ul>\n<p>2. <strong>Optional Components</strong>:</p>\n<ul>\n<li>RateLimiter: Smart request pacing and backoff</li>\n<li>CrawlerMonitor: Real-time progress visualization</li>\n</ul>\n<p>3. <strong>Key Benefits</strong>:</p>\n<ul>\n<li>Automatic memory management</li>\n<li>Built-in rate limiting</li>\n<li>Live progress monitoring</li>\n<li>Flexible concurrency control</li>\n</ul>\n<p>Choose the dispatcher that best fits your needs:</p>\n<ul>\n<li><strong>MemoryAdaptiveDispatcher</strong>: For large crawls or limited resources</li>\n<li><strong>SemaphoreDispatcher</strong>: For simple, fixed-concurrency scenarios</li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/network-console-capture/",
    "title": "Network & Console Capture - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"network-requests-console-message-capturing\">Network Requests &amp; Console Message Capturing</h1>\n<p>Crawl4AI can capture all network requests and browser console messages during a crawl, which is invaluable for debugging, security analysis, or understanding page behavior.</p>\n<h2 id=\"configuration\">Configuration</h2>\n<p>To enable network and console capturing, use these configuration options:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Enable both network request capture and console message capture</span>\nconfig = CrawlerRunConfig(\n    capture_network_requests=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Capture all network requests and responses</span>\n    capture_console_messages=<span class=\"hljs-literal\">True</span>   <span class=\"hljs-comment\"># Capture all browser console output</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"example-usage\">Example Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Enable both network request capture and console message capture</span>\n    config = CrawlerRunConfig(\n        capture_network_requests=<span class=\"hljs-literal\">True</span>,\n        capture_console_messages=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># Analyze network requests</span>\n            <span class=\"hljs-keyword\">if</span> result.network_requests:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.network_requests)}</span> network events\"</span>)\n\n                <span class=\"hljs-comment\"># Count request types</span>\n                request_count = <span class=\"hljs-built_in\">len</span>([r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request\"</span>])\n                response_count = <span class=\"hljs-built_in\">len</span>([r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"response\"</span>])\n                failed_count = <span class=\"hljs-built_in\">len</span>([r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request_failed\"</span>])\n\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Requests: <span class=\"hljs-subst\">{request_count}</span>, Responses: <span class=\"hljs-subst\">{response_count}</span>, Failed: <span class=\"hljs-subst\">{failed_count}</span>\"</span>)\n\n                <span class=\"hljs-comment\"># Find API calls</span>\n                api_calls = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests \n                            <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request\"</span> <span class=\"hljs-keyword\">and</span> <span class=\"hljs-string\">\"api\"</span> <span class=\"hljs-keyword\">in</span> r.get(<span class=\"hljs-string\">\"url\"</span>, <span class=\"hljs-string\">\"\"</span>)]\n                <span class=\"hljs-keyword\">if</span> api_calls:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Detected <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(api_calls)}</span> API calls:\"</span>)\n                    <span class=\"hljs-keyword\">for</span> call <span class=\"hljs-keyword\">in</span> api_calls[:<span class=\"hljs-number\">3</span>]:  <span class=\"hljs-comment\"># Show first 3</span>\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  - <span class=\"hljs-subst\">{call.get(<span class=\"hljs-string\">'method'</span>)}</span> <span class=\"hljs-subst\">{call.get(<span class=\"hljs-string\">'url'</span>)}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Analyze console messages</span>\n            <span class=\"hljs-keyword\">if</span> result.console_messages:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.console_messages)}</span> console messages\"</span>)\n\n                <span class=\"hljs-comment\"># Group by type</span>\n                message_types = {}\n                <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n                    msg_type = msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)\n                    message_types[msg_type] = message_types.get(msg_type, <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Message types:\"</span>, message_types)\n\n                <span class=\"hljs-comment\"># Show errors (often the most important)</span>\n                errors = [msg <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages <span class=\"hljs-keyword\">if</span> msg.get(<span class=\"hljs-string\">\"type\"</span>) == <span class=\"hljs-string\">\"error\"</span>]\n                <span class=\"hljs-keyword\">if</span> errors:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(errors)}</span> console errors:\"</span>)\n                    <span class=\"hljs-keyword\">for</span> err <span class=\"hljs-keyword\">in</span> errors[:<span class=\"hljs-number\">2</span>]:  <span class=\"hljs-comment\"># Show first 2</span>\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  - <span class=\"hljs-subst\">{err.get(<span class=\"hljs-string\">'text'</span>, <span class=\"hljs-string\">''</span>)[:<span class=\"hljs-number\">100</span>]}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Export all captured data to a file for detailed analysis</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"network_capture.json\"</span>, <span class=\"hljs-string\">\"w\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                json.dump({\n                    <span class=\"hljs-string\">\"url\"</span>: result.url,\n                    <span class=\"hljs-string\">\"network_requests\"</span>: result.network_requests <span class=\"hljs-keyword\">or</span> [],\n                    <span class=\"hljs-string\">\"console_messages\"</span>: result.console_messages <span class=\"hljs-keyword\">or</span> []\n                }, f, indent=<span class=\"hljs-number\">2</span>)\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Exported detailed capture data to network_capture.json\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"captured-data-structure\">Captured Data Structure</h2>\n<h3 id=\"network-requests\">Network Requests</h3>\n<p>The <code>result.network_requests</code> contains a list of dictionaries, each representing a network event with these common fields:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Field</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>event_type</code></td>\n<td>Type of event: <code>\"request\"</code>, <code>\"response\"</code>, or <code>\"request_failed\"</code></td>\n</tr>\n<tr>\n<td><code>url</code></td>\n<td>The URL of the request</td>\n</tr>\n<tr>\n<td><code>timestamp</code></td>\n<td>Unix timestamp when the event was captured</td>\n</tr>\n</tbody>\n</table>\n<h4 id=\"request-event-fields\">Request Event Fields</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"event_type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"request\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"url\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"https://example.com/api/data.json\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"method\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"GET\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"headers\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span><span class=\"hljs-attr\">\"User-Agent\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span><span class=\"hljs-punctuation\">,</span> <span class=\"hljs-attr\">\"Accept\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span><span class=\"hljs-punctuation\">}</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"post_data\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"key=value&amp;otherkey=value\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"resource_type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"fetch\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"is_navigation_request\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\"><span class=\"hljs-keyword\">false</span></span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"timestamp\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1633456789.123</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"response-event-fields\">Response Event Fields</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"event_type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"response\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"url\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"https://example.com/api/data.json\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"status\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">200</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"status_text\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"OK\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"headers\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span><span class=\"hljs-attr\">\"Content-Type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"application/json\"</span><span class=\"hljs-punctuation\">,</span> <span class=\"hljs-attr\">\"Cache-Control\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span><span class=\"hljs-punctuation\">}</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"from_service_worker\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\"><span class=\"hljs-keyword\">false</span></span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"request_timing\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span><span class=\"hljs-attr\">\"requestTime\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1234.56</span><span class=\"hljs-punctuation\">,</span> <span class=\"hljs-attr\">\"receiveHeadersEnd\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1234.78</span><span class=\"hljs-punctuation\">}</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"timestamp\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1633456789.456</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"failed-request-event-fields\">Failed Request Event Fields</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"event_type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"request_failed\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"url\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"https://example.com/missing.png\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"method\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"GET\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"resource_type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"image\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"failure_text\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"net::ERR_ABORTED 404\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"timestamp\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1633456789.789</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"console-messages\">Console Messages</h3>\n<p>The <code>result.console_messages</code> contains a list of dictionaries, each representing a console message with these common fields:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Field</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>type</code></td>\n<td>Message type: <code>\"log\"</code>, <code>\"error\"</code>, <code>\"warning\"</code>, <code>\"info\"</code>, etc.</td>\n</tr>\n<tr>\n<td><code>text</code></td>\n<td>The message text</td>\n</tr>\n<tr>\n<td><code>timestamp</code></td>\n<td>Unix timestamp when the message was captured</td>\n</tr>\n</tbody>\n</table>\n<h4 id=\"console-message-example\">Console Message Example</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"error\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"text\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Uncaught TypeError: Cannot read property 'length' of undefined\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"location\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"https://example.com/script.js:123:45\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"timestamp\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1633456790.123</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"key-benefits\">Key Benefits</h2>\n<ul>\n<li><strong>Full Request Visibility</strong>: Capture all network activity including:</li>\n<li>Requests (URLs, methods, headers, post data)</li>\n<li>Responses (status codes, headers, timing)</li>\n<li>\n<p>Failed requests (with error messages)</p>\n</li>\n<li>\n<p><strong>Console Message Access</strong>: View all JavaScript console output:</p>\n</li>\n<li>Log messages</li>\n<li>Warnings</li>\n<li>Errors with stack traces</li>\n<li>\n<p>Developer debugging information</p>\n</li>\n<li>\n<p><strong>Debugging Power</strong>: Identify issues such as:</p>\n</li>\n<li>Failed API calls or resource loading</li>\n<li>JavaScript errors affecting page functionality</li>\n<li>CORS or other security issues</li>\n<li>\n<p>Hidden API endpoints and data flows</p>\n</li>\n<li>\n<p><strong>Security Analysis</strong>: Detect:</p>\n</li>\n<li>Unexpected third-party requests</li>\n<li>Data leakage in request payloads</li>\n<li>\n<p>Suspicious script behavior</p>\n</li>\n<li>\n<p><strong>Performance Insights</strong>: Analyze:</p>\n</li>\n<li>Request timing data</li>\n<li>Resource loading patterns</li>\n<li>Potential bottlenecks</li>\n</ul>\n<h2 id=\"use-cases\">Use Cases</h2>\n<ol>\n<li><strong>API Discovery</strong>: Identify hidden endpoints and data flows in single-page applications</li>\n<li><strong>Debugging</strong>: Track down JavaScript errors affecting page functionality</li>\n<li><strong>Security Auditing</strong>: Detect unwanted third-party requests or data leakage</li>\n<li><strong>Performance Analysis</strong>: Identify slow-loading resources</li>\n<li><strong>Ad/Tracker Analysis</strong>: Detect and catalog advertising or tracking calls</li>\n</ol>\n<p>This capability is especially valuable for complex sites with heavy JavaScript, single-page applications, or when you need to understand the exact communication happening between a browser and servers.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/pdf-parsing/",
    "title": "PDF Parsing - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"pdf-processing-strategies\">PDF Processing Strategies</h1>\n<p>Crawl4AI provides specialized strategies for handling and extracting content from PDF files. These strategies allow you to seamlessly integrate PDF processing into your crawling workflows, whether the PDFs are hosted online or stored locally.</p>\n<h2 id=\"pdfcrawlerstrategy\"><code>PDFCrawlerStrategy</code></h2>\n<h3 id=\"overview\">Overview</h3>\n<p><code>PDFCrawlerStrategy</code> is an implementation of <code>AsyncCrawlerStrategy</code> designed specifically for PDF documents. Instead of interpreting the input URL as an HTML webpage, this strategy treats it as a pointer to a PDF file. It doesn't perform deep crawling or HTML parsing itself but rather prepares the PDF source for a dedicated PDF scraping strategy. Its primary role is to identify the PDF source (web URL or local file) and pass it along the processing pipeline in a way that <code>AsyncWebCrawler</code> can handle.</p>\n<h3 id=\"when-to-use\">When to Use</h3>\n<p>Use <code>PDFCrawlerStrategy</code> when you need to:\n- Process PDF files using the <code>AsyncWebCrawler</code>.\n- Handle PDFs from both web URLs (e.g., <code>https://example.com/document.pdf</code>) and local file paths (e.g., <code>file:///path/to/your/document.pdf</code>).\n- Integrate PDF content extraction into a unified <code>CrawlResult</code> object, allowing consistent handling of PDF data alongside web page data.</p>\n<h3 id=\"key-methods-and-their-behavior\">Key Methods and Their Behavior</h3>\n<ul>\n<li><strong><code>__init__(self, logger: AsyncLogger = None)</code></strong>:<ul>\n<li>Initializes the strategy.</li>\n<li><code>logger</code>: An optional <code>AsyncLogger</code> instance (from <code>crawl4ai.async_logger</code>) for logging purposes.</li>\n</ul>\n</li>\n<li><strong><code>async crawl(self, url: str, **kwargs) -&gt; AsyncCrawlResponse</code></strong>:<ul>\n<li>This method is called by the <code>AsyncWebCrawler</code> during the <code>arun</code> process.</li>\n<li>It takes the <code>url</code> (which should point to a PDF) and creates a minimal <code>AsyncCrawlResponse</code>.</li>\n<li>The <code>html</code> attribute of this response is typically empty or a placeholder, as the actual PDF content processing is deferred to the <code>PDFContentScrapingStrategy</code> (or a similar PDF-aware scraping strategy).</li>\n<li>It sets <code>response_headers</code> to indicate \"application/pdf\" and <code>status_code</code> to 200.</li>\n</ul>\n</li>\n<li><strong><code>async close(self)</code></strong>:<ul>\n<li>A method for cleaning up any resources used by the strategy. For <code>PDFCrawlerStrategy</code>, this is usually minimal.</li>\n</ul>\n</li>\n<li><strong><code>async __aenter__(self)</code> / <code>async __aexit__(self, exc_type, exc_val, exc_tb)</code></strong>:<ul>\n<li>Enables asynchronous context management for the strategy, allowing it to be used with <code>async with</code>.</li>\n</ul>\n</li>\n</ul>\n<h3 id=\"example-usage\">Example Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFCrawlerStrategy, PDFContentScrapingStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Initialize the PDF crawler strategy</span>\n    pdf_crawler_strategy = PDFCrawlerStrategy()\n\n    <span class=\"hljs-comment\"># PDFCrawlerStrategy is typically used in conjunction with PDFContentScrapingStrategy</span>\n    <span class=\"hljs-comment\"># The scraping strategy handles the actual PDF content extraction</span>\n    pdf_scraping_strategy = PDFContentScrapingStrategy()\n    run_config = CrawlerRunConfig(scraping_strategy=pdf_scraping_strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(crawler_strategy=pdf_crawler_strategy) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Example with a remote PDF URL</span>\n        pdf_url = <span class=\"hljs-string\">\"https://arxiv.org/pdf/2310.06825.pdf\"</span> <span class=\"hljs-comment\"># A public PDF from arXiv</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Attempting to process PDF: <span class=\"hljs-subst\">{pdf_url}</span>\"</span>)\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=pdf_url, config=run_config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully processed PDF: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Metadata Title: <span class=\"hljs-subst\">{result.metadata.get(<span class=\"hljs-string\">'title'</span>, <span class=\"hljs-string\">'N/A'</span>)}</span>\"</span>)\n            <span class=\"hljs-comment\"># Further processing of result.markdown, result.media, etc.</span>\n            <span class=\"hljs-comment\"># would be done here, based on what PDFContentScrapingStrategy extracts.</span>\n            <span class=\"hljs-keyword\">if</span> result.markdown <span class=\"hljs-keyword\">and</span> <span class=\"hljs-built_in\">hasattr</span>(result.markdown, <span class=\"hljs-string\">'raw_markdown'</span>):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted text (first 200 chars): <span class=\"hljs-subst\">{result.markdown.raw_markdown[:<span class=\"hljs-number\">200</span>]}</span>...\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"No markdown (text) content extracted.\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to process PDF: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"pros-and-cons\">Pros and Cons</h3>\n<p><strong>Pros:</strong>\n-   Enables <code>AsyncWebCrawler</code> to handle PDF sources directly using familiar <code>arun</code> calls.\n-   Provides a consistent interface for specifying PDF sources (URLs or local paths).\n-   Abstracts the source handling, allowing a separate scraping strategy to focus on PDF content parsing.</p>\n<p><strong>Cons:</strong>\n-   Does not perform any PDF data extraction itself; it strictly relies on a compatible scraping strategy (like <code>PDFContentScrapingStrategy</code>) to process the PDF.\n-   Has limited utility on its own; most of its value comes from being paired with a PDF-specific content scraping strategy.</p>\n<hr>\n<h2 id=\"pdfcontentscrapingstrategy\"><code>PDFContentScrapingStrategy</code></h2>\n<h3 id=\"overview_1\">Overview</h3>\n<p><code>PDFContentScrapingStrategy</code> is an implementation of <code>ContentScrapingStrategy</code> designed to extract text, metadata, and optionally images from PDF documents. It is intended to be used in conjunction with a crawler strategy that can provide it with a PDF source, such as <code>PDFCrawlerStrategy</code>. This strategy uses the <code>NaivePDFProcessorStrategy</code> internally to perform the low-level PDF parsing.</p>\n<h3 id=\"when-to-use_1\">When to Use</h3>\n<p>Use <code>PDFContentScrapingStrategy</code> when your <code>AsyncWebCrawler</code> (often configured with <code>PDFCrawlerStrategy</code>) needs to:\n-   Extract textual content page by page from a PDF document.\n-   Retrieve standard metadata embedded within the PDF (e.g., title, author, subject, creation date, page count).\n-   Optionally, extract images contained within the PDF pages. These images can be saved to a local directory or made available for further processing.\n-   Produce a <code>ScrapingResult</code> that can be converted into a <code>CrawlResult</code>, making PDF content accessible in a manner similar to HTML web content (e.g., text in <code>result.markdown</code>, metadata in <code>result.metadata</code>).</p>\n<h3 id=\"key-configuration-attributes\">Key Configuration Attributes</h3>\n<p>When initializing <code>PDFContentScrapingStrategy</code>, you can configure its behavior using the following attributes:\n-   <strong><code>extract_images: bool = False</code></strong>: If <code>True</code>, the strategy will attempt to extract images from the PDF.\n-   <strong><code>save_images_locally: bool = False</code></strong>: If <code>True</code> (and <code>extract_images</code> is also <code>True</code>), extracted images will be saved to disk in the <code>image_save_dir</code>. If <code>False</code>, image data might be available in another form (e.g., base64, depending on the underlying processor) but not saved as separate files by this strategy.\n-   <strong><code>image_save_dir: str = None</code></strong>: Specifies the directory where extracted images should be saved if <code>save_images_locally</code> is <code>True</code>. If <code>None</code>, a default or temporary directory might be used.\n-   <strong><code>batch_size: int = 4</code></strong>: Defines how many PDF pages are processed in a single batch. This can be useful for managing memory when dealing with very large PDF documents.\n-   <strong><code>logger: AsyncLogger = None</code></strong>: An optional <code>AsyncLogger</code> instance for logging.</p>\n<h3 id=\"key-methods-and-their-behavior_1\">Key Methods and Their Behavior</h3>\n<ul>\n<li><strong><code>__init__(self, save_images_locally: bool = False, extract_images: bool = False, image_save_dir: str = None, batch_size: int = 4, logger: AsyncLogger = None)</code></strong>:<ul>\n<li>Initializes the strategy with configurations for image handling, batch processing, and logging. It sets up an internal <code>NaivePDFProcessorStrategy</code> instance which performs the actual PDF parsing.</li>\n</ul>\n</li>\n<li><strong><code>scrap(self, url: str, html: str, **params) -&gt; ScrapingResult</code></strong>:<ul>\n<li>This is the primary synchronous method called by the crawler (via <code>ascrap</code>) to process the PDF.</li>\n<li><code>url</code>: The path or URL to the PDF file (provided by <code>PDFCrawlerStrategy</code> or similar).</li>\n<li><code>html</code>: Typically an empty string when used with <code>PDFCrawlerStrategy</code>, as the content is a PDF, not HTML.</li>\n<li>It first ensures the PDF is accessible locally (downloads it to a temporary file if <code>url</code> is remote).</li>\n<li>It then uses its internal PDF processor to extract text, metadata, and images (if configured).</li>\n<li>The extracted information is compiled into a <code>ScrapingResult</code> object:<ul>\n<li><code>cleaned_html</code>: Contains an HTML-like representation of the PDF, where each page's content is often wrapped in a <code>&lt;div&gt;</code> with page number information.</li>\n<li><code>media</code>: A dictionary where <code>media[\"images\"]</code> will contain information about extracted images if <code>extract_images</code> was <code>True</code>.</li>\n<li><code>links</code>: A dictionary where <code>links[\"urls\"]</code> can contain URLs found within the PDF content.</li>\n<li><code>metadata</code>: A dictionary holding PDF metadata (e.g., title, author, num_pages).</li>\n</ul>\n</li>\n</ul>\n</li>\n<li><strong><code>async ascrap(self, url: str, html: str, **kwargs) -&gt; ScrapingResult</code></strong>:<ul>\n<li>The asynchronous version of <code>scrap</code>. Under the hood, it typically runs the synchronous <code>scrap</code> method in a separate thread using <code>asyncio.to_thread</code> to avoid blocking the event loop.</li>\n</ul>\n</li>\n<li><strong><code>_get_pdf_path(self, url: str) -&gt; str</code></strong>:<ul>\n<li>A private helper method to manage PDF file access. If the <code>url</code> is remote (http/https), it downloads the PDF to a temporary local file and returns its path. If <code>url</code> indicates a local file (<code>file://</code> or a direct path), it resolves and returns the local path.</li>\n</ul>\n</li>\n</ul>\n<h3 id=\"example-usage_1\">Example Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFCrawlerStrategy, PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">import</span> os <span class=\"hljs-comment\"># For creating image directory</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Define the directory for saving extracted images</span>\n    image_output_dir = <span class=\"hljs-string\">\"./my_pdf_images\"</span>\n    os.makedirs(image_output_dir, exist_ok=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-comment\"># Configure the PDF content scraping strategy</span>\n    <span class=\"hljs-comment\"># Enable image extraction and specify where to save them</span>\n    pdf_scraping_cfg = PDFContentScrapingStrategy(\n        extract_images=<span class=\"hljs-literal\">True</span>,\n        save_images_locally=<span class=\"hljs-literal\">True</span>,\n        image_save_dir=image_output_dir,\n        batch_size=<span class=\"hljs-number\">2</span> <span class=\"hljs-comment\"># Process 2 pages at a time for demonstration</span>\n    )\n\n    <span class=\"hljs-comment\"># The PDFCrawlerStrategy is needed to tell AsyncWebCrawler how to \"crawl\" a PDF</span>\n    pdf_crawler_cfg = PDFCrawlerStrategy()\n\n    <span class=\"hljs-comment\"># Configure the overall crawl run</span>\n    run_cfg = CrawlerRunConfig(\n        scraping_strategy=pdf_scraping_cfg <span class=\"hljs-comment\"># Use our PDF scraping strategy</span>\n    )\n\n    <span class=\"hljs-comment\"># Initialize the crawler with the PDF-specific crawler strategy</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(crawler_strategy=pdf_crawler_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        pdf_url = <span class=\"hljs-string\">\"https://arxiv.org/pdf/2310.06825.pdf\"</span> <span class=\"hljs-comment\"># Example PDF</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Starting PDF processing for: <span class=\"hljs-subst\">{pdf_url}</span>\"</span>)\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=pdf_url, config=run_cfg)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n--- PDF Processing Successful ---\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Processed URL: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n--- Metadata ---\"</span>)\n            <span class=\"hljs-keyword\">for</span> key, value <span class=\"hljs-keyword\">in</span> result.metadata.items():\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  <span class=\"hljs-subst\">{key.replace(<span class=\"hljs-string\">'_'</span>, <span class=\"hljs-string\">' '</span>).title()}</span>: <span class=\"hljs-subst\">{value}</span>\"</span>)\n\n            <span class=\"hljs-keyword\">if</span> result.markdown <span class=\"hljs-keyword\">and</span> <span class=\"hljs-built_in\">hasattr</span>(result.markdown, <span class=\"hljs-string\">'raw_markdown'</span>):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n--- Extracted Text (Markdown Snippet) ---\"</span>)\n                <span class=\"hljs-built_in\">print</span>(result.markdown.raw_markdown[:<span class=\"hljs-number\">500</span>].strip() + <span class=\"hljs-string\">\"...\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nNo text (markdown) content extracted.\"</span>)\n\n            <span class=\"hljs-keyword\">if</span> result.media <span class=\"hljs-keyword\">and</span> result.media.get(<span class=\"hljs-string\">\"images\"</span>):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n--- Image Extraction ---\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.media[<span class=\"hljs-string\">'images'</span>])}</span> image(s).\"</span>)\n                <span class=\"hljs-keyword\">for</span> i, img_info <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(result.media[<span class=\"hljs-string\">\"images\"</span>][:<span class=\"hljs-number\">2</span>]): <span class=\"hljs-comment\"># Show info for first 2 images</span>\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Image <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>:\"</span>)\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"    Page: <span class=\"hljs-subst\">{img_info.get(<span class=\"hljs-string\">'page'</span>)}</span>\"</span>)\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"    Format: <span class=\"hljs-subst\">{img_info.get(<span class=\"hljs-string\">'format'</span>, <span class=\"hljs-string\">'N/A'</span>)}</span>\"</span>)\n                    <span class=\"hljs-keyword\">if</span> img_info.get(<span class=\"hljs-string\">'path'</span>):\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"    Saved at: <span class=\"hljs-subst\">{img_info.get(<span class=\"hljs-string\">'path'</span>)}</span>\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nNo images were extracted (or extract_images was False).\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n--- PDF Processing Failed ---\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"pros-and-cons_1\">Pros and Cons</h3>\n<p><strong>Pros:</strong>\n-   Provides a comprehensive way to extract text, metadata, and (optionally) images from PDF documents.\n-   Handles both remote PDFs (via URL) and local PDF files.\n-   Configurable image extraction allows saving images to disk or accessing their data.\n-   Integrates smoothly with the <code>CrawlResult</code> object structure, making PDF-derived data accessible in a way consistent with web-scraped data.\n-   The <code>batch_size</code> parameter can help in managing memory consumption when processing large or numerous PDF pages.</p>\n<p><strong>Cons:</strong>\n-   Extraction quality and performance can vary significantly depending on the PDF's complexity, encoding, and whether it's image-based (scanned) or text-based.\n-   Image extraction can be resource-intensive (both CPU and disk space if <code>save_images_locally</code> is true).\n-   Relies on <code>NaivePDFProcessorStrategy</code> internally, which might have limitations with very complex layouts, encrypted PDFs, or forms compared to more sophisticated PDF parsing libraries. Scanned PDFs will not yield text unless an OCR step is performed (which is not part of this strategy by default).\n-   Link extraction from PDFs can be basic and depends on how hyperlinks are embedded in the document.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/proxy-security/",
    "title": "Proxy & Security - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"proxy-security\">Proxy &amp; Security</h1>\n<p>This guide covers proxy configuration and security features in Crawl4AI, including SSL certificate analysis and proxy rotation strategies.</p>\n<h2 id=\"understanding-proxy-configuration\">Understanding Proxy Configuration</h2>\n<p>Crawl4AI recommends configuring proxies per request through <code>CrawlerRunConfig.proxy_config</code>. This gives you precise control, enables rotation strategies, and keeps examples simple enough to copy, paste, and run.</p>\n<h2 id=\"basic-proxy-setup\">Basic Proxy Setup</h2>\n<p>Configure proxies that apply to each crawl operation:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, ProxyConfig\n\nrun_config = CrawlerRunConfig(proxy_config=ProxyConfig(server=<span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>))\n<span class=\"hljs-comment\"># run_config = CrawlerRunConfig(proxy_config={\"server\": \"http://proxy.example.com:8080\"})</span>\n<span class=\"hljs-comment\"># run_config = CrawlerRunConfig(proxy_config=\"http://proxy.example.com:8080\")</span>\n\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig()\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Success: <span class=\"hljs-subst\">{result.success}</span> -&gt; <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"admonition note\">\n<p class=\"admonition-title\">Why request-level?</p>\n<p><code>CrawlerRunConfig.proxy_config</code> keeps each request self-contained, so swapping proxies or rotation strategies is just a matter of building a new run configuration.</p>\n</div>\n<h2 id=\"supported-proxy-formats\">Supported Proxy Formats</h2>\n<p>The <code>ProxyConfig.from_string()</code> method supports multiple formats:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai import ProxyConfig\n\n<span class=\"hljs-comment\"># HTTP proxy with authentication</span>\nproxy1 = ProxyConfig.from_string(<span class=\"hljs-string\">\"http://user:pass@192.168.1.1:8080\"</span>)\n\n<span class=\"hljs-comment\"># HTTPS proxy</span>\nproxy2 = ProxyConfig.from_string(<span class=\"hljs-string\">\"https://proxy.example.com:8080\"</span>)\n\n<span class=\"hljs-comment\"># SOCKS5 proxy</span>\nproxy3 = ProxyConfig.from_string(<span class=\"hljs-string\">\"socks5://proxy.example.com:1080\"</span>)\n\n<span class=\"hljs-comment\"># Simple IP:port format</span>\nproxy4 = ProxyConfig.from_string(<span class=\"hljs-string\">\"192.168.1.1:8080\"</span>)\n\n<span class=\"hljs-comment\"># IP:port:user:pass format</span>\nproxy5 = ProxyConfig.from_string(<span class=\"hljs-string\">\"192.168.1.1:8080:user:pass\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"authenticated-proxies\">Authenticated Proxies</h2>\n<p>For proxies requiring authentication:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler,BrowserConfig, CrawlerRunConfig, ProxyConfig\n\nrun_config = CrawlerRunConfig(\n    proxy_config=ProxyConfig(\n        server=<span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>,\n        username=<span class=\"hljs-string\">\"your_username\"</span>,\n        password=<span class=\"hljs-string\">\"your_password\"</span>,\n    )\n)\n<span class=\"hljs-comment\"># Or dictionary style:</span>\n<span class=\"hljs-comment\"># run_config = CrawlerRunConfig(proxy_config={</span>\n<span class=\"hljs-comment\">#     \"server\": \"http://proxy.example.com:8080\",</span>\n<span class=\"hljs-comment\">#     \"username\": \"your_username\",</span>\n<span class=\"hljs-comment\">#     \"password\": \"your_password\",</span>\n<span class=\"hljs-comment\"># })</span>\n\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig()\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Success: <span class=\"hljs-subst\">{result.success}</span> -&gt; <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"environment-variable-configuration\">Environment Variable Configuration</h2>\n<p>Load proxies from environment variables for easy configuration:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> ProxyConfig, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Set environment variable</span>\nos.environ[<span class=\"hljs-string\">\"PROXIES\"</span>] = <span class=\"hljs-string\">\"ip1:port1:user1:pass1,ip2:port2:user2:pass2,ip3:port3\"</span>\n\n<span class=\"hljs-comment\"># Load all proxies</span>\nproxies = ProxyConfig.from_env()\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Loaded <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(proxies)}</span> proxies\"</span>)\n\n<span class=\"hljs-comment\"># Use first proxy</span>\n<span class=\"hljs-keyword\">if</span> proxies:\n    run_config = CrawlerRunConfig(proxy_config=proxies[<span class=\"hljs-number\">0</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"rotating-proxies\">Rotating Proxies</h2>\n<p>Crawl4AI supports automatic proxy rotation to distribute requests across multiple proxy servers. Rotation is applied per request using a rotation strategy on <code>CrawlerRunConfig</code>.</p>\n<h3 id=\"proxy-rotation-recommended\">Proxy Rotation (recommended)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> re\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, ProxyConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.proxy_strategy <span class=\"hljs-keyword\">import</span> RoundRobinProxyStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Load proxies from environment</span>\n    proxies = ProxyConfig.from_env()\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> proxies:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"No proxies found! Set PROXIES environment variable.\"</span>)\n        <span class=\"hljs-keyword\">return</span>\n\n    <span class=\"hljs-comment\"># Create rotation strategy</span>\n    proxy_strategy = RoundRobinProxyStrategy(proxies)\n\n    <span class=\"hljs-comment\"># Configure per-request with proxy rotation</span>\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, verbose=<span class=\"hljs-literal\">False</span>)\n    run_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        proxy_rotation_strategy=proxy_strategy,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        urls = [<span class=\"hljs-string\">\"https://httpbin.org/ip\"</span>] * (<span class=\"hljs-built_in\">len</span>(proxies) * <span class=\"hljs-number\">2</span>)  <span class=\"hljs-comment\"># Test each proxy twice</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🚀 Testing <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(proxies)}</span> proxies with rotation...\"</span>)\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls=urls, config=run_config)\n\n        <span class=\"hljs-keyword\">for</span> i, result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(results):\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-comment\"># Extract IP from response</span>\n                ip_match = re.search(<span class=\"hljs-string\">r'(?:[0-9]{1,3}\\.){3}[0-9]{1,3}'</span>, result.html)\n                <span class=\"hljs-keyword\">if</span> ip_match:\n                    detected_ip = ip_match.group(<span class=\"hljs-number\">0</span>)\n                    proxy_index = i % <span class=\"hljs-built_in\">len</span>(proxies)\n                    expected_ip = proxies[proxy_index].ip\n\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✅ Request <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>: Proxy <span class=\"hljs-subst\">{proxy_index+<span class=\"hljs-number\">1</span>}</span> -&gt; IP <span class=\"hljs-subst\">{detected_ip}</span>\"</span>)\n                    <span class=\"hljs-keyword\">if</span> detected_ip == expected_ip:\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"   🎯 IP matches proxy configuration\"</span>)\n                    <span class=\"hljs-keyword\">else</span>:\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   ⚠️  IP mismatch (expected <span class=\"hljs-subst\">{expected_ip}</span>)\"</span>)\n                <span class=\"hljs-keyword\">else</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"❌ Request <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>: Could not extract IP from response\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"❌ Request <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>: Failed - <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"ssl-certificate-analysis\">SSL Certificate Analysis</h2>\n<p>Combine proxy usage with SSL certificate inspection for enhanced security analysis. SSL certificate fetching is configured per request via <code>CrawlerRunConfig</code>.</p>\n<h3 id=\"per-request-ssl-certificate-analysis\">Per-Request SSL Certificate Analysis</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\nrun_config = CrawlerRunConfig(\n    proxy_config={\n        <span class=\"hljs-string\">\"server\"</span>: <span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>,\n        <span class=\"hljs-string\">\"username\"</span>: <span class=\"hljs-string\">\"user\"</span>,\n        <span class=\"hljs-string\">\"password\"</span>: <span class=\"hljs-string\">\"pass\"</span>,\n    },\n    fetch_ssl_certificate=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable SSL certificate analysis for this request</span>\n)\n\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig()\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✅ Crawled via proxy: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Analyze SSL certificate</span>\n            <span class=\"hljs-keyword\">if</span> result.ssl_certificate:\n                cert = result.ssl_certificate\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"🔒 SSL Certificate Info:\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Issuer: <span class=\"hljs-subst\">{cert.issuer}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Subject: <span class=\"hljs-subst\">{cert.subject}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Valid until: <span class=\"hljs-subst\">{cert.valid_until}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Fingerprint: <span class=\"hljs-subst\">{cert.fingerprint}</span>\"</span>)\n\n                <span class=\"hljs-comment\"># Export certificate</span>\n                cert.to_json(<span class=\"hljs-string\">\"certificate.json\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"💾 Certificate exported to certificate.json\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"⚠️  No SSL certificate information available\"</span>)\n\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"security-best-practices\">Security Best Practices</h2>\n<h3 id=\"1-proxy-rotation-for-anonymity\">1. Proxy Rotation for Anonymity</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai import CrawlerRunConfig, ProxyConfig\nfrom crawl4ai.proxy_strategy import RoundRobinProxyStrategy\n\n<span class=\"hljs-comment\"># Use multiple proxies to avoid IP blocking</span>\nproxies = ProxyConfig.from_env(<span class=\"hljs-string\">\"PROXIES\"</span>)\nstrategy = RoundRobinProxyStrategy(proxies)\n\n<span class=\"hljs-comment\"># Configure rotation per request (recommended)</span>\nrun_config = CrawlerRunConfig(proxy_rotation_strategy=strategy)\n\n<span class=\"hljs-comment\"># For a fixed proxy across all requests, just reuse the same run_config instance</span>\nstatic_run_config = run_config\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-ssl-certificate-verification\">2. SSL Certificate Verification</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Always verify SSL certificates when possible</span>\n<span class=\"hljs-comment\"># Per-request (affects specific requests)</span>\nrun_config = CrawlerRunConfig(fetch_ssl_certificate=<span class=\"hljs-literal\">True</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-environment-variable-security\">3. Environment Variable Security</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Use environment variables for sensitive proxy credentials</span>\n<span class=\"hljs-comment\"># Avoid hardcoding usernames/passwords in code</span>\n<span class=\"hljs-built_in\">export</span> PROXIES=<span class=\"hljs-string\">\"ip1:port1:user1:pass1,ip2:port2:user2:pass2\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"4-socks5-for-enhanced-security\">4. SOCKS5 for Enhanced Security</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Prefer SOCKS5 proxies for better protocol support</span>\nrun_config = CrawlerRunConfig(proxy_config=<span class=\"hljs-string\">\"socks5://proxy.example.com:1080\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"migration-from-deprecated-proxy-parameter\">Migration from Deprecated <code>proxy</code> Parameter</h2>\n<ul>\n<li>\"Deprecation Notice\"\n    The legacy <code>proxy</code> argument on <code>BrowserConfig</code> is deprecated. Configure proxies through <code>CrawlerRunConfig.proxy_config</code> so each request fully describes its network settings.</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Old (deprecated) approach</span>\n<span class=\"hljs-comment\"># from crawl4ai import BrowserConfig</span>\n<span class=\"hljs-comment\"># browser_config = BrowserConfig(proxy_config=\"http://proxy.example.com:8080\")</span>\n\n<span class=\"hljs-comment\"># New (preferred) approach</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\nrun_config = CrawlerRunConfig(proxy_config=<span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"safe-logging-of-proxies\">Safe Logging of Proxies</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> ProxyConfig\n\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">safe_proxy_repr</span>(<span class=\"hljs-params\">proxy: ProxyConfig</span>):\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">getattr</span>(proxy, <span class=\"hljs-string\">\"username\"</span>, <span class=\"hljs-literal\">None</span>):\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{proxy.server}</span> (auth: ****)\"</span>\n    <span class=\"hljs-keyword\">return</span> proxy.server\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"troubleshooting\">Troubleshooting</h2>\n<h3 id=\"common-issues\">Common Issues</h3>\n<ul>\n<li>\n<p>\"Proxy connection failed\"</p>\n<ul>\n<li>Verify the proxy server is reachable from your network.</li>\n<li>Double-check authentication credentials.</li>\n<li>Ensure the protocol matches (<code>http</code>, <code>https</code>, or <code>socks5</code>).</li>\n</ul>\n</li>\n<li>\n<p>\"SSL certificate errors\"</p>\n<ul>\n<li>Some proxies break SSL inspection; switch proxies if you see repeated failures.</li>\n<li>Consider temporarily disabling certificate fetching to isolate the issue.</li>\n</ul>\n</li>\n<li>\n<p>\"Environment variables not loading\"</p>\n<ul>\n<li>Confirm <code>PROXIES</code> (or your custom env var) is set before running the script.</li>\n<li>Check formatting: <code>ip:port:user:pass,ip:port:user:pass</code>.</li>\n</ul>\n</li>\n<li>\n<p>\"Proxy rotation not working\"</p>\n<ul>\n<li>Ensure <code>ProxyConfig.from_env()</code> actually loaded entries (<code>len(proxies) &gt; 0</code>).</li>\n<li>Attach <code>proxy_rotation_strategy</code> to <code>CrawlerRunConfig</code>.</li>\n<li>Validate the proxy definitions you pass into the strategy.</li>\n</ul>\n</li>\n</ul>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../anti-bot-and-fallback/\">Anti-Bot Detection &amp; Fallback</a> — Automatic retry with proxy escalation and fallback functions when anti-bot blocking is detected</li>\n</ul>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/session-management/",
    "title": "Session Management - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"session-management\">Session Management</h1>\n<p>Session management in Crawl4AI is a powerful feature that allows you to maintain state across multiple requests, making it particularly suitable for handling complex multi-step crawling tasks. It enables you to reuse the same browser tab (or page object) across sequential actions and crawls, which is beneficial for:</p>\n<ul>\n<li><strong>Performing JavaScript actions before and after crawling.</strong></li>\n<li><strong>Executing multiple sequential crawls faster</strong> without needing to reopen tabs or allocate memory repeatedly.</li>\n</ul>\n<p><strong>Note:</strong> This feature is designed for sequential workflows and is not suitable for parallel operations.</p>\n<hr>\n<h4 id=\"basic-session-usage\">Basic Session Usage</h4>\n<p>Use <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> to maintain state with a <code>session_id</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs import BrowserConfig, <span class=\"hljs-function\">CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    session_id</span> = <span class=\"hljs-string\">\"my_session\"</span>\n\n    <span class=\"hljs-meta\"># Define configurations</span>\n    config1 = CrawlerRunConfig(\n        url=<span class=\"hljs-string\">\"https://example.com/page1\"</span>, session_id=session_id\n    )\n    config2 = CrawlerRunConfig(\n        url=<span class=\"hljs-string\">\"https://example.com/page2\"</span>, session_id=session_id\n    )\n\n    <span class=\"hljs-meta\"># First request</span>\n    result1 = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config1)\n\n    <span class=\"hljs-meta\"># Subsequent request using the same session</span>\n    result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config2)\n\n    <span class=\"hljs-meta\"># Clean up when done</span>\n    <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h4 id=\"dynamic-content-with-sessions\">Dynamic Content with Sessions</h4>\n<p>Here's an example of crawling GitHub commits across multiple pages while preserving session state:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.cache_context <span class=\"hljs-keyword\">import</span> CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_dynamic_content</span>():\n    url = <span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>\n    session_id = <span class=\"hljs-string\">\"wait_for_session\"</span>\n    all_commits = []\n\n    js_next_page = <span class=\"hljs-string\">\"\"\"\n    const commits = document.querySelectorAll('li[data-testid=\"commit-row-item\"] h4');\n    if (commits.length &gt; 0) {\n        window.lastCommit = commits[0].textContent.trim();\n    }\n    const button = document.querySelector('a[data-testid=\"pagination-next-button\"]');\n    if (button) {button.click(); console.log('button clicked') }\n    \"\"\"</span>\n\n    wait_for = <span class=\"hljs-string\">\"\"\"() =&gt; {\n        const commits = document.querySelectorAll('li[data-testid=\"commit-row-item\"] h4');\n        if (commits.length === 0) return false;\n        const firstCommit = commits[0].textContent.trim();\n        return firstCommit !== window.lastCommit;\n    }\"\"\"</span>\n\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Commit Extractor\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"li[data-testid='commit-row-item']\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h4 a\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n                <span class=\"hljs-string\">\"transform\"</span>: <span class=\"hljs-string\">\"strip\"</span>,\n            },\n        ],\n    }\n    extraction_strategy = JsonCssExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n\n    browser_config = BrowserConfig(\n        verbose=<span class=\"hljs-literal\">True</span>,\n        headless=<span class=\"hljs-literal\">False</span>,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            crawler_config = CrawlerRunConfig(\n                session_id=session_id,\n                css_selector=<span class=\"hljs-string\">\"li[data-testid='commit-row-item']\"</span>,\n                extraction_strategy=extraction_strategy,\n                js_code=js_next_page <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                wait_for=wait_for <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS,\n                capture_console_messages=<span class=\"hljs-literal\">True</span>,\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=crawler_config)\n\n            <span class=\"hljs-keyword\">if</span> result.console_messages:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span> console messages:\"</span>, result.console_messages)\n\n            <span class=\"hljs-keyword\">if</span> result.extracted_content:\n                <span class=\"hljs-comment\"># print(f\"Page {page + 1} result:\", result.extracted_content)</span>\n                commits = json.loads(result.extracted_content)\n                all_commits.extend(commits)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(commits)}</span> commits\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: No content extracted\"</span>)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(all_commits)}</span> commits across 3 pages\"</span>)\n        <span class=\"hljs-comment\"># Clean up session</span>\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"example-1-basic-session-based-crawling\">Example 1: Basic Session-Based Crawling</h2>\n<p>A simple example using session-based crawling:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.cache_context <span class=\"hljs-keyword\">import</span> CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">basic_session_crawl</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"dynamic_content_session\"</span>\n        url = <span class=\"hljs-string\">\"https://example.com/dynamic-content\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=<span class=\"hljs-string\">\"document.querySelector('.load-more-button').click();\"</span> <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\".content-item\"</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{result.extracted_content.count(<span class=\"hljs-string\">'.content-item'</span>)}</span> items\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(basic_session_crawl())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This example shows:\n1. Reusing the same <code>session_id</code> across multiple requests.\n2. Executing JavaScript to load more content dynamically.\n3. Properly closing the session to free resources.</p>\n<hr>\n<h2 id=\"advanced-technique-1-custom-execution-hooks\">Advanced Technique 1: Custom Execution Hooks</h2>\n<blockquote>\n<p>Warning: You might feel confused by the end of the next few examples 😅, so make sure you are comfortable with the order of the parts before you start this.</p>\n</blockquote>\n<p>Use custom hooks to handle complex scenarios, such as waiting for content to load dynamically:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">advanced_session_crawl_with_hooks</span>():\n    first_commit = <span class=\"hljs-string\">\"\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_execution_started</span>(<span class=\"hljs-params\">page</span>):\n        <span class=\"hljs-keyword\">nonlocal</span> first_commit\n        <span class=\"hljs-keyword\">try</span>:\n            <span class=\"hljs-keyword\">while</span> <span class=\"hljs-literal\">True</span>:\n                <span class=\"hljs-keyword\">await</span> page.wait_for_selector(<span class=\"hljs-string\">\"li.commit-item h4\"</span>)\n                commit = <span class=\"hljs-keyword\">await</span> page.query_selector(<span class=\"hljs-string\">\"li.commit-item h4\"</span>)\n                commit = <span class=\"hljs-keyword\">await</span> commit.evaluate(<span class=\"hljs-string\">\"(element) =&gt; element.textContent\"</span>).strip()\n                <span class=\"hljs-keyword\">if</span> commit <span class=\"hljs-keyword\">and</span> commit != first_commit:\n                    first_commit = commit\n                    <span class=\"hljs-keyword\">break</span>\n                <span class=\"hljs-keyword\">await</span> asyncio.sleep(<span class=\"hljs-number\">0.5</span>)\n        <span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Warning: New content didn't appear: <span class=\"hljs-subst\">{e}</span>\"</span>)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"commit_session\"</span>\n        url = <span class=\"hljs-string\">\"https://github.com/example/repo/commits/main\"</span>\n        crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"on_execution_started\"</span>, on_execution_started)\n\n        js_next_page = <span class=\"hljs-string\">\"\"\"document.querySelector('a.pagination-next').click();\"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=js_next_page <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\"li.commit-item\"</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.extracted_content)}</span> commits\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(advanced_session_crawl_with_hooks())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This technique ensures new content loads before the next action.</p>\n<hr>\n<h2 id=\"advanced-technique-2-integrated-javascript-execution-and-waiting\">Advanced Technique 2: Integrated JavaScript Execution and Waiting</h2>\n<p>Combine JavaScript execution and waiting logic for concise handling of dynamic content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">integrated_js_and_wait_crawl</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"integrated_session\"</span>\n        url = <span class=\"hljs-string\">\"https://github.com/example/repo/commits/main\"</span>\n\n        js_next_page_and_wait = <span class=\"hljs-string\">\"\"\"\n        (async () =&gt; {\n            const getCurrentCommit = () =&gt; document.querySelector('li.commit-item h4').textContent.trim();\n            const initialCommit = getCurrentCommit();\n            document.querySelector('a.pagination-next').click();\n            while (getCurrentCommit() === initialCommit) {\n                await new Promise(resolve =&gt; setTimeout(resolve, 100));\n            }\n        })();\n        \"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=js_next_page_and_wait <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\"li.commit-item\"</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.extracted_content)}</span> commits\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(integrated_js_and_wait_crawl())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h4 id=\"common-use-cases-for-sessions\">Common Use Cases for Sessions</h4>\n<p>1. <strong>Authentication Flows</strong>: Login and interact with secured pages.</p>\n<p>2. <strong>Pagination Handling</strong>: Navigate through multiple pages.</p>\n<p>3. <strong>Form Submissions</strong>: Fill forms, submit, and process results.</p>\n<p>4. <strong>Multi-step Processes</strong>: Complete workflows that span multiple actions.</p>\n<p>5. <strong>Dynamic Content Navigation</strong>: Handle JavaScript-rendered or event-triggered content.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/ssl-certificate/",
    "title": "SSL Certificate - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"sslcertificate-reference\"><code>SSLCertificate</code> Reference</h1>\n<p>The <strong><code>SSLCertificate</code></strong> class encapsulates an SSL certificate’s data and allows exporting it in various formats (PEM, DER, JSON, or text). It’s used within <strong>Crawl4AI</strong> whenever you set <strong><code>fetch_ssl_certificate=True</code></strong> in your <strong><code>CrawlerRunConfig</code></strong>.  </p>\n<h2 id=\"1-overview\">1. Overview</h2>\n<p><strong>Location</strong>: <code>crawl4ai/ssl_certificate.py</code></p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">SSLCertificate</span>:\n    <span class=\"hljs-string\">\"\"\"\n    Represents an SSL certificate with methods to export in various formats.\n\n    Main Methods:\n    - from_url(url, timeout=10)\n    - from_file(file_path)\n    - from_binary(binary_data)\n    - to_json(filepath=None)\n    - to_pem(filepath=None)\n    - to_der(filepath=None)\n    ...\n\n    Common Properties:\n    - issuer\n    - subject\n    - valid_from\n    - valid_until\n    - fingerprint\n    \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"typical-use-case\">Typical Use Case</h3>\n<ol>\n<li>You <strong>enable</strong> certificate fetching in your crawl by:\n   <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\"><span class=\"hljs-built_in\">CrawlerRunConfig</span>(fetch_ssl_certificate=True, ...)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li>After <code>arun()</code>, if <code>result.ssl_certificate</code> is present, it’s an instance of <strong><code>SSLCertificate</code></strong>.  </li>\n<li>You can <strong>read</strong> basic properties (issuer, subject, validity) or <strong>export</strong> them in multiple formats.</li>\n</ol>\n<hr>\n<h2 id=\"2-construction-fetching\">2. Construction &amp; Fetching</h2>\n<h3 id=\"21-from_urlurl-timeout10\">2.1 <strong><code>from_url(url, timeout=10)</code></strong></h3>\n<p>Manually load an SSL certificate from a given URL (port 443). Typically used internally, but you can call it directly if you want:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">cert = SSLCertificate.from_url(<span class=\"hljs-string\">\"https://example.com\"</span>)\n<span class=\"hljs-keyword\">if</span> cert:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fingerprint:\"</span>, cert.fingerprint)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"22-from_filefile_path\">2.2 <strong><code>from_file(file_path)</code></strong></h3>\n<p>Load from a file containing certificate data in ASN.1 or DER. Rarely needed unless you have local cert files:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">cert</span> = SSLCertificate.from_file(<span class=\"hljs-string\">\"/path/to/cert.der\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"23-from_binarybinary_data\">2.3 <strong><code>from_binary(binary_data)</code></strong></h3>\n<p>Initialize from raw binary. E.g., if you captured it from a socket or another source:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">cert</span> = SSLCertificate.from_binary(raw_bytes)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"3-common-properties\">3. Common Properties</h2>\n<p>After obtaining a <strong><code>SSLCertificate</code></strong> instance (e.g. <code>result.ssl_certificate</code> from a crawl), you can read:</p>\n<p>1. <strong><code>issuer</code></strong> <em>(dict)</em><br>\n   - E.g. <code>{\"CN\": \"My Root CA\", \"O\": \"...\"}</code>\n2. <strong><code>subject</code></strong> <em>(dict)</em><br>\n   - E.g. <code>{\"CN\": \"example.com\", \"O\": \"ExampleOrg\"}</code>\n3. <strong><code>valid_from</code></strong> <em>(str)</em><br>\n   - NotBefore date/time. Often in ASN.1/UTC format.\n4. <strong><code>valid_until</code></strong> <em>(str)</em><br>\n   - NotAfter date/time.\n5. <strong><code>fingerprint</code></strong> <em>(str)</em><br>\n   - The SHA-256 digest (lowercase hex).<br>\n   - E.g. <code>\"d14d2e...\"</code></p>\n<hr>\n<h2 id=\"4-export-methods\">4. Export Methods</h2>\n<p>Once you have a <strong><code>SSLCertificate</code></strong> object, you can <strong>export</strong> or <strong>inspect</strong> it:</p>\n<h3 id=\"41-to_jsonfilepathnone-optionalstr\">4.1 <strong><code>to_json(filepath=None)</code> → <code>Optional[str]</code></strong></h3>\n<ul>\n<li>Returns a JSON string containing the parsed certificate fields.  </li>\n<li>If <code>filepath</code> is provided, saves it to disk instead, returning <code>None</code>.</li>\n</ul>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">json_data = cert.to_json()  <span class=\"hljs-comment\"># returns JSON string</span>\ncert.to_json(<span class=\"hljs-string\">\"certificate.json\"</span>)  <span class=\"hljs-comment\"># writes file, returns None</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"42-to_pemfilepathnone-optionalstr\">4.2 <strong><code>to_pem(filepath=None)</code> → <code>Optional[str]</code></strong></h3>\n<ul>\n<li>Returns a PEM-encoded string (common for web servers).  </li>\n<li>If <code>filepath</code> is provided, saves it to disk instead.</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">pem_str = cert.to_pem()              <span class=\"hljs-comment\"># in-memory PEM string</span>\ncert.to_pem(<span class=\"hljs-string\">\"/path/to/cert.pem\"</span>)     <span class=\"hljs-comment\"># saved to file</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"43-to_derfilepathnone-optionalbytes\">4.3 <strong><code>to_der(filepath=None)</code> → <code>Optional[bytes]</code></strong></h3>\n<ul>\n<li>Returns the original DER (binary ASN.1) bytes.  </li>\n<li>If <code>filepath</code> is specified, writes the bytes there instead.</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">der_bytes = cert.to_der()\ncert.to_der(<span class=\"hljs-string\">\"certificate.der\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"44-optional-export_as_text\">4.4 (Optional) <strong><code>export_as_text()</code></strong></h3>\n<ul>\n<li>If you see a method like <code>export_as_text()</code>, it typically returns an OpenSSL-style textual representation.  </li>\n<li>Not always needed, but can help for debugging or manual inspection.</li>\n</ul>\n<hr>\n<h2 id=\"5-example-usage-in-crawl4ai\">5. Example Usage in Crawl4AI</h2>\n<p>Below is a minimal sample showing how the crawler obtains an SSL cert from a site, then reads or exports it. The code snippet:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">import asyncio\nimport os\n<span class=\"hljs-selector-tag\">from</span> crawl4ai import AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\nasync def <span class=\"hljs-selector-tag\">main</span>():\n    tmp_dir = <span class=\"hljs-string\">\"tmp\"</span>\n    os.<span class=\"hljs-built_in\">makedirs</span>(tmp_dir, exist_ok=True)\n\n    config = <span class=\"hljs-built_in\">CrawlerRunConfig</span>(\n        fetch_ssl_certificate=True,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    async with <span class=\"hljs-built_in\">AsyncWebCrawler</span>() as crawler:\n        result = await crawler.<span class=\"hljs-built_in\">arun</span>(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n        if result.success and result.ssl_certificate:\n            cert = result.ssl_certificate\n            # <span class=\"hljs-number\">1</span>. Basic Info\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Issuer CN:\"</span>, cert.issuer.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"CN\"</span>, <span class=\"hljs-string\">\"\"</span>))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Valid until:\"</span>, cert.valid_until)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fingerprint:\"</span>, cert.fingerprint)\n\n            # <span class=\"hljs-number\">2</span>. Export\n            cert.<span class=\"hljs-built_in\">to_json</span>(os.path.<span class=\"hljs-built_in\">join</span>(tmp_dir, <span class=\"hljs-string\">\"certificate.json\"</span>))\n            cert.<span class=\"hljs-built_in\">to_pem</span>(os.path.<span class=\"hljs-built_in\">join</span>(tmp_dir, <span class=\"hljs-string\">\"certificate.pem\"</span>))\n            cert.<span class=\"hljs-built_in\">to_der</span>(os.path.<span class=\"hljs-built_in\">join</span>(tmp_dir, <span class=\"hljs-string\">\"certificate.der\"</span>))\n\nif __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.<span class=\"hljs-built_in\">run</span>(<span class=\"hljs-built_in\">main</span>())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"6-notes-best-practices\">6. Notes &amp; Best Practices</h2>\n<p>1. <strong>Timeout</strong>: <code>SSLCertificate.from_url</code> internally uses a default <strong>10s</strong> socket connect and wraps SSL.<br>\n2. <strong>Binary Form</strong>: The certificate is loaded in ASN.1 (DER) form, then re-parsed by <code>OpenSSL.crypto</code>.<br>\n3. <strong>Validation</strong>: This does <strong>not</strong> validate the certificate chain or trust store. It only fetches and parses.<br>\n4. <strong>Integration</strong>: Within Crawl4AI, you typically just set <code>fetch_ssl_certificate=True</code> in <code>CrawlerRunConfig</code>; the final result’s <code>ssl_certificate</code> is automatically built.<br>\n5. <strong>Export</strong>: If you need to store or analyze a cert, the <code>to_json</code> and <code>to_pem</code> are quite universal.</p>\n<hr>\n<h3 id=\"summary\">Summary</h3>\n<ul>\n<li><strong><code>SSLCertificate</code></strong> is a convenience class for capturing and exporting the <strong>TLS certificate</strong> from your crawled site(s).  </li>\n<li>Common usage is in the <strong><code>CrawlResult.ssl_certificate</code></strong> field, accessible after setting <code>fetch_ssl_certificate=True</code>.  </li>\n<li>Offers quick access to essential certificate details (<code>issuer</code>, <code>subject</code>, <code>fingerprint</code>) and is easy to export (PEM, DER, JSON) for further analysis or server usage.</li>\n</ul>\n<p>Use it whenever you need <strong>insight</strong> into a site’s certificate or require some form of cryptographic or compliance check.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/undetected-browser/",
    "title": "Undetected Browser - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"undetected-browser-mode\">Undetected Browser Mode</h1>\n<h2 id=\"overview\">Overview</h2>\n<p>Crawl4AI offers two powerful anti-bot features to help you access websites with bot detection:</p>\n<ol>\n<li><strong>Stealth Mode</strong> - Uses playwright-stealth to modify browser fingerprints and behaviors</li>\n<li><strong>Undetected Browser Mode</strong> - Advanced browser adapter with deep-level patches for sophisticated bot detection</li>\n</ol>\n<p>This guide covers both features and helps you choose the right approach for your needs.</p>\n<h2 id=\"anti-bot-features-comparison\">Anti-Bot Features Comparison</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th>Regular Browser</th>\n<th>Stealth Mode</th>\n<th>Undetected Browser</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>WebDriver Detection</td>\n<td>❌</td>\n<td>✅</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>Navigator Properties</td>\n<td>❌</td>\n<td>✅</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>Plugin Emulation</td>\n<td>❌</td>\n<td>✅</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>CDP Detection</td>\n<td>❌</td>\n<td>Partial</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>Deep Browser Patches</td>\n<td>❌</td>\n<td>❌</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>Performance Impact</td>\n<td>None</td>\n<td>Minimal</td>\n<td>Moderate</td>\n</tr>\n<tr>\n<td>Setup Complexity</td>\n<td>None</td>\n<td>None</td>\n<td>Minimal</td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"when-to-use-each-approach\">When to Use Each Approach</h2>\n<h3 id=\"use-regular-browser-stealth-mode-when\">Use Regular Browser + Stealth Mode When:</h3>\n<ul>\n<li>Sites have basic bot detection (checking navigator.webdriver, plugins, etc.)</li>\n<li>You need good performance with basic protection</li>\n<li>Sites check for common automation indicators</li>\n</ul>\n<h3 id=\"use-undetected-browser-when\">Use Undetected Browser When:</h3>\n<ul>\n<li>Sites employ sophisticated bot detection services (Cloudflare, DataDome, etc.)</li>\n<li>Stealth mode alone isn't sufficient</li>\n<li>You're willing to trade some performance for better evasion</li>\n</ul>\n<h3 id=\"best-practice-progressive-enhancement\">Best Practice: Progressive Enhancement</h3>\n<ol>\n<li><strong>Start with</strong>: Regular browser + Stealth mode</li>\n<li><strong>If blocked</strong>: Switch to Undetected browser</li>\n<li><strong>If still blocked</strong>: Combine Undetected browser + Stealth mode</li>\n</ol>\n<h2 id=\"stealth-mode\">Stealth Mode</h2>\n<p>Stealth mode is the simpler anti-bot solution that works with both regular and undetected browsers:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\n<span class=\"hljs-comment\"># Enable stealth mode with regular browser</span>\nbrowser_config = BrowserConfig(\n    enable_stealth=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Simple flag to enable</span>\n    headless=<span class=\"hljs-literal\">False</span>       <span class=\"hljs-comment\"># Better for avoiding detection</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"what-stealth-mode-does\">What Stealth Mode Does:</h3>\n<ul>\n<li>Removes <code>navigator.webdriver</code> flag</li>\n<li>Modifies browser fingerprints</li>\n<li>Emulates realistic plugin behavior</li>\n<li>Adjusts navigator properties</li>\n<li>Fixes common automation leaks</li>\n</ul>\n<h2 id=\"undetected-browser-mode_1\">Undetected Browser Mode</h2>\n<p>For sites with sophisticated bot detection that stealth mode can't bypass, use the undetected browser adapter:</p>\n<h3 id=\"key-features\">Key Features</h3>\n<ul>\n<li><strong>Drop-in Replacement</strong>: Uses the same API as regular browser mode</li>\n<li><strong>Enhanced Stealth</strong>: Built-in patches to evade common detection methods</li>\n<li><strong>Browser Adapter Pattern</strong>: Seamlessly switch between regular and undetected modes</li>\n<li><strong>Automatic Installation</strong>: <code>crawl4ai-setup</code> installs all necessary browser dependencies</li>\n</ul>\n<h3 id=\"quick-start\">Quick Start</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler, \n    BrowserConfig, \n    CrawlerRunConfig,\n    UndetectedAdapter\n)\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_crawler_strategy <span class=\"hljs-keyword\">import</span> AsyncPlaywrightCrawlerStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Create the undetected adapter</span>\n    undetected_adapter = UndetectedAdapter()\n\n    <span class=\"hljs-comment\"># Create browser config</span>\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Headless mode can be detected easier</span>\n        verbose=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-comment\"># Create the crawler strategy with undetected adapter</span>\n    crawler_strategy = AsyncPlaywrightCrawlerStrategy(\n        browser_config=browser_config,\n        browser_adapter=undetected_adapter\n    )\n\n    <span class=\"hljs-comment\"># Create the crawler with our custom strategy</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(\n        crawler_strategy=crawler_strategy,\n        config=browser_config\n    ) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Your crawling code here</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=CrawlerRunConfig()\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">500</span>])\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"combining-both-features\">Combining Both Features</h2>\n<p>For maximum evasion, combine stealth mode with undetected browser:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, UndetectedAdapter\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_crawler_strategy <span class=\"hljs-keyword\">import</span> AsyncPlaywrightCrawlerStrategy\n\n<span class=\"hljs-comment\"># Create browser config with stealth enabled</span>\nbrowser_config = BrowserConfig(\n    enable_stealth=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable stealth mode</span>\n    headless=<span class=\"hljs-literal\">False</span>\n)\n\n<span class=\"hljs-comment\"># Create undetected adapter</span>\nadapter = UndetectedAdapter()\n\n<span class=\"hljs-comment\"># Create strategy with both features</span>\nstrategy = AsyncPlaywrightCrawlerStrategy(\n    browser_config=browser_config,\n    browser_adapter=adapter\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(\n    crawler_strategy=strategy,\n    config=browser_config\n) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://protected-site.com\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"examples\">Examples</h2>\n<h3 id=\"example-1-basic-stealth-mode\">Example 1: Basic Stealth Mode</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">test_stealth_mode</span>():\n    <span class=\"hljs-comment\"># Simple stealth mode configuration</span>\n    browser_config = BrowserConfig(\n        enable_stealth=<span class=\"hljs-literal\">True</span>,\n        headless=<span class=\"hljs-literal\">False</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://bot.sannysoft.com\"</span>,\n            config=CrawlerRunConfig(screenshot=<span class=\"hljs-literal\">True</span>)\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"✓ Successfully accessed bot detection test site\"</span>)\n            <span class=\"hljs-comment\"># Save screenshot to verify detection results</span>\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-keyword\">import</span> base64\n                <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"stealth_test.png\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                    f.write(base64.b64decode(result.screenshot))\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"✓ Screenshot saved - check for green (passed) tests\"</span>)\n\nasyncio.run(test_stealth_mode())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"example-2-undetected-browser-mode\">Example 2: Undetected Browser Mode</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    BrowserConfig,\n    CrawlerRunConfig,\n    UndetectedAdapter\n)\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_crawler_strategy <span class=\"hljs-keyword\">import</span> AsyncPlaywrightCrawlerStrategy\n\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Create browser config</span>\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,\n        verbose=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-comment\"># Create the undetected adapter</span>\n    undetected_adapter = UndetectedAdapter()\n\n    <span class=\"hljs-comment\"># Create the crawler strategy with the undetected adapter</span>\n    crawler_strategy = AsyncPlaywrightCrawlerStrategy(\n        browser_config=browser_config,\n        browser_adapter=undetected_adapter\n    )\n\n    <span class=\"hljs-comment\"># Create the crawler with our custom strategy</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(\n        crawler_strategy=crawler_strategy,\n        config=browser_config\n    ) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Configure the crawl</span>\n        crawler_config = CrawlerRunConfig(\n            markdown_generator=DefaultMarkdownGenerator(\n                content_filter=PruningContentFilter()\n            ),\n            capture_console_messages=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Test adapter console capture</span>\n        )\n\n        <span class=\"hljs-comment\"># Test on a site that typically detects bots</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Testing undetected adapter...\"</span>)\n        result: CrawlResult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.helloworld.org\"</span>, \n            config=crawler_config\n        )\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Status: <span class=\"hljs-subst\">{result.status_code}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Success: <span class=\"hljs-subst\">{result.success}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Console messages captured: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.console_messages <span class=\"hljs-keyword\">or</span> [])}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Markdown content (first 500 chars):\\n<span class=\"hljs-subst\">{result.markdown.raw_markdown[:<span class=\"hljs-number\">500</span>]}</span>\"</span>)\n\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"browser-adapter-pattern\">Browser Adapter Pattern</h2>\n<p>The undetected browser support is implemented using an adapter pattern, allowing seamless switching between different browser implementations:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Regular browser adapter (default)</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> PlaywrightAdapter\nregular_adapter = PlaywrightAdapter()\n\n<span class=\"hljs-comment\"># Undetected browser adapter</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> UndetectedAdapter\nundetected_adapter = UndetectedAdapter()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The adapter handles:\n- JavaScript execution\n- Console message capture\n- Error handling\n- Browser-specific optimizations</p>\n<h2 id=\"best-practices\">Best Practices</h2>\n<ol>\n<li>\n<p><strong>Avoid Headless Mode</strong>: Detection is easier in headless mode\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">browser_config</span> = BrowserConfig(headless=<span class=\"hljs-literal\">False</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Use Reasonable Delays</strong>: Don't rush through pages\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">crawler_config = CrawlerRunConfig(\n    wait_time=3.0,  <span class=\"hljs-comment\"># Wait 3 seconds after page load</span>\n    delay_before_return_html=2.0  <span class=\"hljs-comment\"># Additional delay</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Rotate User Agents</strong>: You can customize user agents\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">browser_config = BrowserConfig(\n    headers={<span class=\"hljs-string\">\"User-Agent\"</span>: <span class=\"hljs-string\">\"your-user-agent\"</span>}\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Handle Failures Gracefully</strong>: Some sites may still detect and block\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n<h2 id=\"advanced-usage-tips\">Advanced Usage Tips</h2>\n<h3 id=\"progressive-detection-handling\">Progressive Detection Handling</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_with_progressive_evasion</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-comment\"># Step 1: Try regular browser with stealth</span>\n    browser_config = BrowserConfig(\n        enable_stealth=<span class=\"hljs-literal\">True</span>,\n        headless=<span class=\"hljs-literal\">False</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url)\n        <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> <span class=\"hljs-string\">\"Access Denied\"</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-keyword\">in</span> result.html:\n            <span class=\"hljs-keyword\">return</span> result\n\n    <span class=\"hljs-comment\"># Step 2: If blocked, try undetected browser</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Regular + stealth blocked, trying undetected browser...\"</span>)\n\n    adapter = UndetectedAdapter()\n    strategy = AsyncPlaywrightCrawlerStrategy(\n        browser_config=browser_config,\n        browser_adapter=adapter\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(\n        crawler_strategy=strategy,\n        config=browser_config\n    ) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url)\n        <span class=\"hljs-keyword\">return</span> result\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"installation\">Installation</h2>\n<p>The undetected browser dependencies are automatically installed when you run:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This command installs all necessary browser dependencies for both regular and undetected modes.</p>\n<h2 id=\"limitations\">Limitations</h2>\n<ul>\n<li><strong>Performance</strong>: Slightly slower than regular mode due to additional patches</li>\n<li><strong>Headless Detection</strong>: Some sites can still detect headless mode</li>\n<li><strong>Resource Usage</strong>: May use more resources than regular mode</li>\n<li><strong>Not 100% Guaranteed</strong>: Advanced anti-bot services are constantly evolving</li>\n</ul>\n<h2 id=\"troubleshooting\">Troubleshooting</h2>\n<h3 id=\"browser-not-found\">Browser Not Found</h3>\n<p>Run the setup command:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"detection-still-occurring\">Detection Still Occurring</h3>\n<p>Try combining with other features:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    simulate_user<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Add user simulation</span>\n    magic<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable magic mode</span>\n    wait_time<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5.0</span>,  <span class=\"hljs-comment\"># Longer waits</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"performance-issues\">Performance Issues</h3>\n<p>If experiencing slow performance:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Use selective undetected mode only for protected sites</span>\n<span class=\"hljs-keyword\">if</span> is_protected_site(url):\n    adapter = UndetectedAdapter()\n<span class=\"hljs-keyword\">else</span>:\n    adapter = PlaywrightAdapter()  <span class=\"hljs-comment\"># Default adapter</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"future-plans\">Future Plans</h2>\n<p><strong>Note</strong>: In future versions of Crawl4AI, we may enable stealth mode and undetected browser by default to provide better out-of-the-box success rates. For now, users should explicitly enable these features when needed.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Crawl4AI provides flexible anti-bot solutions:</p>\n<ol>\n<li><strong>Start Simple</strong>: Use regular browser + stealth mode for most sites</li>\n<li><strong>Escalate if Needed</strong>: Switch to undetected browser for sophisticated protection</li>\n<li><strong>Combine for Maximum Effect</strong>: Use both features together when facing the toughest challenges</li>\n</ol>\n<p>Remember:\n- Always respect robots.txt and website terms of service\n- Use appropriate delays to avoid overwhelming servers\n- Consider the performance trade-offs of each approach\n- Test progressively to find the minimum necessary evasion level</p>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../advanced-features/\">Advanced Features</a> - Overview of all advanced features</li>\n<li><a href=\"../proxy-security/\">Proxy &amp; Security</a> - Using proxies with anti-bot features</li>\n<li><a href=\"../session-management/\">Session Management</a> - Maintaining sessions across requests</li>\n<li><a href=\"../identity-based-crawling/\">Identity Based Crawling</a> - Additional anti-detection strategies</li>\n<li><a href=\"../anti-bot-and-fallback/\">Anti-Bot Detection &amp; Fallback</a> - Automatic retry and proxy escalation when blocking is detected</li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/advanced/virtual-scroll/",
    "title": "Virtual Scroll - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"virtual-scroll\">Virtual Scroll</h1>\n<p>Modern websites increasingly use <strong>virtual scrolling</strong> (also called windowed rendering or viewport rendering) to handle large datasets efficiently. This technique only renders visible items in the DOM, replacing content as users scroll. Popular examples include Twitter's timeline, Instagram's feed, and many data tables.</p>\n<p>Crawl4AI's Virtual Scroll feature automatically detects and handles these scenarios, ensuring you capture <strong>all content</strong>, not just what's initially visible.</p>\n<h2 id=\"understanding-virtual-scroll\">Understanding Virtual Scroll</h2>\n<h3 id=\"the-problem\">The Problem</h3>\n<p>Traditional infinite scroll <strong>appends</strong> new content to existing content. Virtual scroll <strong>replaces</strong> content to maintain performance:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\">Traditional <span class=\"hljs-keyword\">Scroll</span>:          Virtual <span class=\"hljs-keyword\">Scroll</span>:\n┌─────────────┐             ┌─────────────┐\n│ Item <span class=\"hljs-number\">1</span>      │             │ Item <span class=\"hljs-number\">11</span>     │  <span class=\"hljs-operator\">&lt;</span><span class=\"hljs-operator\">-</span> Items <span class=\"hljs-number\">1</span><span class=\"hljs-number\">-10</span> removed\n│ Item <span class=\"hljs-number\">2</span>      │             │ Item <span class=\"hljs-number\">12</span>     │  <span class=\"hljs-operator\">&lt;</span><span class=\"hljs-operator\">-</span> <span class=\"hljs-keyword\">Only</span> visible items\n│ ...         │             │ Item <span class=\"hljs-number\">13</span>     │     <span class=\"hljs-keyword\">in</span> DOM\n│ Item <span class=\"hljs-number\">10</span>     │             │ Item <span class=\"hljs-number\">14</span>     │\n│ Item <span class=\"hljs-number\">11</span> <span class=\"hljs-keyword\">NEW</span> │             │ Item <span class=\"hljs-number\">15</span>     │\n│ Item <span class=\"hljs-number\">12</span> <span class=\"hljs-keyword\">NEW</span> │             └─────────────┘\n└─────────────┘             \nDOM keeps growing           DOM size stays constant\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Without proper handling, crawlers only capture the currently visible items, missing the rest of the content.</p>\n<h3 id=\"three-scrolling-scenarios\">Three Scrolling Scenarios</h3>\n<p>Crawl4AI's Virtual Scroll detects and handles three scenarios:</p>\n<ol>\n<li><strong>No Change</strong> - Content doesn't update on scroll (static page or end reached)</li>\n<li><strong>Content Appended</strong> - New items added to existing ones (traditional infinite scroll)  </li>\n<li><strong>Content Replaced</strong> - Items replaced with new ones (true virtual scroll)</li>\n</ol>\n<p>Only scenario 3 requires special handling, which Virtual Scroll automates.</p>\n<h2 id=\"basic-usage\">Basic Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, VirtualScrollConfig\n\n<span class=\"hljs-comment\"># Configure virtual scroll</span>\nvirtual_config = VirtualScrollConfig(\n    container_selector=<span class=\"hljs-string\">\"#feed\"</span>,      <span class=\"hljs-comment\"># CSS selector for scrollable container</span>\n    scroll_count=<span class=\"hljs-number\">20</span>,                 <span class=\"hljs-comment\"># Number of scrolls to perform</span>\n    scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,    <span class=\"hljs-comment\"># How much to scroll each time</span>\n    wait_after_scroll=<span class=\"hljs-number\">0.5</span>           <span class=\"hljs-comment\"># Wait time (seconds) after each scroll</span>\n)\n\n<span class=\"hljs-comment\"># Use in crawler configuration</span>\nconfig = CrawlerRunConfig(\n    virtual_scroll_config=virtual_config\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n    <span class=\"hljs-comment\"># result.html contains ALL items from the virtual scroll</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"configuration-parameters\">Configuration Parameters</h2>\n<h3 id=\"virtualscrollconfig\">VirtualScrollConfig</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Type</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>container_selector</code></td>\n<td><code>str</code></td>\n<td>Required</td>\n<td>CSS selector for the scrollable container</td>\n</tr>\n<tr>\n<td><code>scroll_count</code></td>\n<td><code>int</code></td>\n<td><code>10</code></td>\n<td>Maximum number of scrolls to perform</td>\n</tr>\n<tr>\n<td><code>scroll_by</code></td>\n<td><code>str</code> or <code>int</code></td>\n<td><code>\"container_height\"</code></td>\n<td>Scroll amount per step</td>\n</tr>\n<tr>\n<td><code>wait_after_scroll</code></td>\n<td><code>float</code></td>\n<td><code>0.5</code></td>\n<td>Seconds to wait after each scroll</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"scroll-by-options\">Scroll By Options</h3>\n<ul>\n<li><code>\"container_height\"</code> - Scroll by the container's visible height</li>\n<li><code>\"page_height\"</code> - Scroll by the viewport height</li>\n<li><code>500</code> (integer) - Scroll by exact pixel amount</li>\n</ul>\n<h2 id=\"real-world-examples\">Real-World Examples</h2>\n<h3 id=\"twitter-like-timeline\">Twitter-like Timeline</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, VirtualScrollConfig, BrowserConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_twitter_timeline</span>():\n    <span class=\"hljs-comment\"># Twitter replaces tweets as you scroll</span>\n    virtual_config = VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\"[data-testid='primaryColumn']\"</span>,\n        scroll_count=<span class=\"hljs-number\">30</span>,\n        scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,\n        wait_after_scroll=<span class=\"hljs-number\">1.0</span>  <span class=\"hljs-comment\"># Twitter needs time to load</span>\n    )\n\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)  <span class=\"hljs-comment\"># Set to False to watch it work</span>\n    config = CrawlerRunConfig(\n        virtual_scroll_config=virtual_config\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://twitter.com/search?q=AI\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-comment\"># Extract tweet count</span>\n        <span class=\"hljs-keyword\">import</span> re\n        tweets = re.findall(<span class=\"hljs-string\">r'data-testid=\"tweet\"'</span>, result.html)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(tweets)}</span> tweets\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"instagram-grid\">Instagram Grid</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_instagram_grid</span>():\n    <span class=\"hljs-comment\"># Instagram uses virtualized grid for performance</span>\n    virtual_config = VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\"article\"</span>,  <span class=\"hljs-comment\"># Main feed container</span>\n        scroll_count=<span class=\"hljs-number\">50</span>,               <span class=\"hljs-comment\"># More scrolls for grid layout</span>\n        scroll_by=<span class=\"hljs-number\">800</span>,                 <span class=\"hljs-comment\"># Fixed pixel scrolling</span>\n        wait_after_scroll=<span class=\"hljs-number\">0.8</span>\n    )\n\n    config = CrawlerRunConfig(\n        virtual_scroll_config=virtual_config,\n        screenshot=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Capture final state</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.instagram.com/explore/tags/photography/\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-comment\"># Count posts</span>\n        posts = result.html.count(<span class=\"hljs-string\">'class=\"post\"'</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{posts}</span> posts from virtualized grid\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"mixed-content-news-feed\">Mixed Content (News Feed)</h3>\n<p>Some sites mix static and virtualized content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_mixed_feed</span>():\n    <span class=\"hljs-comment\"># Featured articles stay, regular articles virtualize</span>\n    virtual_config = VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\".main-feed\"</span>,\n        scroll_count=<span class=\"hljs-number\">25</span>,\n        scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,\n        wait_after_scroll=<span class=\"hljs-number\">0.5</span>\n    )\n\n    config = CrawlerRunConfig(\n        virtual_scroll_config=virtual_config\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.example.com\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-comment\"># Featured articles remain throughout</span>\n        featured = result.html.count(<span class=\"hljs-string\">'class=\"featured-article\"'</span>)\n        regular = result.html.count(<span class=\"hljs-string\">'class=\"regular-article\"'</span>)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Featured (static): <span class=\"hljs-subst\">{featured}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Regular (virtualized): <span class=\"hljs-subst\">{regular}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"virtual-scroll-vs-scan_full_page\">Virtual Scroll vs scan_full_page</h2>\n<p>Both features handle dynamic content, but serve different purposes:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th>Virtual Scroll</th>\n<th>scan_full_page</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Purpose</strong></td>\n<td>Capture content that's replaced during scroll</td>\n<td>Load content that's appended during scroll</td>\n</tr>\n<tr>\n<td><strong>Use Case</strong></td>\n<td>Twitter, Instagram, virtual tables</td>\n<td>Traditional infinite scroll, lazy-loaded images</td>\n</tr>\n<tr>\n<td><strong>DOM Behavior</strong></td>\n<td>Replaces elements</td>\n<td>Adds elements</td>\n</tr>\n<tr>\n<td><strong>Memory Usage</strong></td>\n<td>Efficient (merges content)</td>\n<td>Can grow large</td>\n</tr>\n<tr>\n<td><strong>Configuration</strong></td>\n<td>Requires container selector</td>\n<td>Works on full page</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"when-to-use-which\">When to Use Which?</h3>\n<p>Use <strong>Virtual Scroll</strong> when:\n- Content disappears as you scroll (Twitter timeline)\n- DOM element count stays relatively constant\n- You need ALL items from a virtualized list\n- Container-based scrolling (not full page)</p>\n<p>Use <strong>scan_full_page</strong> when:\n- Content accumulates as you scroll\n- Images load lazily\n- Simple \"load more\" behavior\n- Full page scrolling</p>\n<h2 id=\"combining-with-extraction\">Combining with Extraction</h2>\n<p>Virtual Scroll works seamlessly with extraction strategies:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy, LLMConfig\n\n<span class=\"hljs-comment\"># Define extraction schema</span>\nschema = {\n    <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"array\"</span>,\n    <span class=\"hljs-string\">\"items\"</span>: {\n        <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"object\"</span>, \n        <span class=\"hljs-string\">\"properties\"</span>: {\n            <span class=\"hljs-string\">\"author\"</span>: {<span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"string\"</span>},\n            <span class=\"hljs-string\">\"content\"</span>: {<span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"string\"</span>},\n            <span class=\"hljs-string\">\"timestamp\"</span>: {<span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"string\"</span>}\n        }\n    }\n}\n\n<span class=\"hljs-comment\"># Configure both virtual scroll and extraction</span>\nconfig = CrawlerRunConfig(\n    virtual_scroll_config=VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\"#timeline\"</span>,\n        scroll_count=<span class=\"hljs-number\">20</span>\n    ),\n    extraction_strategy=LLMExtractionStrategy(\n        llm_config=LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>),\n        schema=schema\n    )\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"...\"</span>, config=config)\n\n    <span class=\"hljs-comment\"># Extracted data from ALL scrolled content</span>\n    <span class=\"hljs-keyword\">import</span> json\n    posts = json.loads(result.extracted_content)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(posts)}</span> posts from virtual scroll\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"performance-tips\">Performance Tips</h2>\n<ol>\n<li>\n<p><strong>Container Selection</strong>: Be specific with selectors. Using the correct container improves performance.</p>\n</li>\n<li>\n<p><strong>Scroll Count</strong>: Start conservative and increase as needed:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Start with fewer scrolls</span>\nvirtual_config = VirtualScrollConfig(\n    container_selector=<span class=\"hljs-string\">\"#feed\"</span>,\n    scroll_count=10  <span class=\"hljs-comment\"># Test with 10, increase if needed</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Wait Times</strong>: Adjust based on site speed:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Fast sites</span>\n<span class=\"hljs-attr\">wait_after_scroll</span>=<span class=\"hljs-number\">0.2</span>\n\n<span class=\"hljs-comment\"># Slower sites or heavy content</span>\n<span class=\"hljs-attr\">wait_after_scroll</span>=<span class=\"hljs-number\">1.5</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Debug Mode</strong>: Set <code>headless=False</code> to watch scrolling:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">False</span>)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Watch the scrolling happen</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n<h2 id=\"how-it-works-internally\">How It Works Internally</h2>\n<ol>\n<li><strong>Detection Phase</strong>: Scrolls and compares HTML to detect behavior</li>\n<li><strong>Capture Phase</strong>: For replaced content, stores HTML chunks at each position</li>\n<li><strong>Merge Phase</strong>: Combines all chunks, removing duplicates based on text content</li>\n<li><strong>Result</strong>: Complete HTML with all unique items</li>\n</ol>\n<p>The deduplication uses normalized text (lowercase, no spaces/symbols) to ensure accurate merging without false positives.</p>\n<h2 id=\"error-handling\">Error Handling</h2>\n<p>Virtual Scroll handles errors gracefully:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># If container not found or scrolling fails</span>\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"...\"</span>, config=config)\n\n<span class=\"hljs-keyword\">if</span> result.success:\n    <span class=\"hljs-comment\"># Virtual scroll worked or wasn't needed</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.html)}</span> characters\"</span>)\n<span class=\"hljs-keyword\">else</span>:\n    <span class=\"hljs-comment\"># Crawl failed entirely</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>If the container isn't found, crawling continues normally without virtual scroll.</p>\n<h2 id=\"complete-example\">Complete Example</h2>\n<p>See our <a href=\"/docs/examples/virtual_scroll_example.py\">comprehensive example</a> that demonstrates:\n- Twitter-like feeds\n- Instagram grids<br>\n- Traditional infinite scroll\n- Mixed content scenarios\n- Performance comparisons</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Run the examples</span>\n<span class=\"hljs-built_in\">cd</span> docs/examples\npython virtual_scroll_example.py\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The example includes a local test server with different scrolling behaviors for experimentation.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/adaptive-crawler/",
    "title": "AdaptiveCrawler - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"adaptivecrawler\">AdaptiveCrawler</h1>\n<p>The <code>AdaptiveCrawler</code> class implements intelligent web crawling that automatically determines when sufficient information has been gathered to answer a query. It uses a three-layer scoring system to evaluate coverage, consistency, and saturation.</p>\n<h2 id=\"constructor\">Constructor</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-less\"><span class=\"hljs-selector-tag\">AdaptiveCrawler</span>(\n    <span class=\"hljs-attribute\">crawler</span>: AsyncWebCrawler,\n    <span class=\"hljs-attribute\">config</span>: Optional[AdaptiveConfig] = None\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"parameters\">Parameters</h3>\n<ul>\n<li><strong>crawler</strong> (<code>AsyncWebCrawler</code>): The underlying web crawler instance to use for fetching pages</li>\n<li><strong>config</strong> (<code>Optional[AdaptiveConfig]</code>): Configuration settings for adaptive crawling behavior. If not provided, uses default settings.</li>\n</ul>\n<h2 id=\"primary-method\">Primary Method</h2>\n<h3 id=\"digest\">digest()</h3>\n<p>The main method that performs adaptive crawling starting from a URL with a specific query.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">digest</span>(<span class=\"hljs-params\">\n    start_url: <span class=\"hljs-built_in\">str</span>,\n    query: <span class=\"hljs-built_in\">str</span>,\n    resume_from: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, Path]] = <span class=\"hljs-literal\">None</span>\n</span>) -&gt; CrawlState\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"parameters_1\">Parameters</h4>\n<ul>\n<li><strong>start_url</strong> (<code>str</code>): The starting URL for crawling</li>\n<li><strong>query</strong> (<code>str</code>): The search query that guides the crawling process</li>\n<li><strong>resume_from</strong> (<code>Optional[Union[str, Path]]</code>): Path to a saved state file to resume from</li>\n</ul>\n<h4 id=\"returns\">Returns</h4>\n<ul>\n<li><strong>CrawlState</strong>: The final crawl state containing all crawled URLs, knowledge base, and metrics</li>\n</ul>\n<h4 id=\"example\">Example</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    adaptive</span> = AdaptiveCrawler(crawler)\n    state = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n        start_url=<span class=\"hljs-string\">\"https://docs.python.org\"</span>,\n        query=<span class=\"hljs-string\">\"async context managers\"</span>\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"properties\">Properties</h2>\n<h3 id=\"confidence\">confidence</h3>\n<p>Current confidence score (0-1) indicating information sufficiency.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-meta\">@property</span>\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">confidence</span>(<span class=\"hljs-params\">self</span>) -&gt; <span class=\"hljs-built_in\">float</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"coverage_stats\">coverage_stats</h3>\n<p>Dictionary containing detailed coverage statistics.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-meta\">@property  </span>\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">coverage_stats</span>(<span class=\"hljs-params\">self</span>) -&gt; <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-built_in\">float</span>]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Returns:\n- <strong>coverage</strong>: Query term coverage score\n- <strong>consistency</strong>: Information consistency score<br>\n- <strong>saturation</strong>: Content saturation score\n- <strong>confidence</strong>: Overall confidence score</p>\n<h3 id=\"is_sufficient\">is_sufficient</h3>\n<p>Boolean indicating whether sufficient information has been gathered.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-meta\">@property</span>\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">is_sufficient</span>(<span class=\"hljs-params\">self</span>) -&gt; <span class=\"hljs-built_in\">bool</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"state\">state</h3>\n<p>Access to the current crawl state.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ruby\"><span class=\"hljs-variable\">@property</span>\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">state</span>(<span class=\"hljs-params\"><span class=\"hljs-variable language_\">self</span></span>) -&gt; <span class=\"hljs-title class_\">CrawlState</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"methods\">Methods</h2>\n<h3 id=\"get_relevant_content\">get_relevant_content()</h3>\n<p>Retrieve the most relevant content from the knowledge base.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">get_relevant_content</span>(<span class=\"hljs-params\">\n    self,\n    top_k: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">5</span>\n</span>) -&gt; <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">Any</span>]]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"parameters_2\">Parameters</h4>\n<ul>\n<li><strong>top_k</strong> (<code>int</code>): Number of top relevant documents to return (default: 5)</li>\n</ul>\n<h4 id=\"returns_1\">Returns</h4>\n<p>List of dictionaries containing:\n- <strong>url</strong>: The URL of the page\n- <strong>content</strong>: The page content\n- <strong>score</strong>: Relevance score\n- <strong>metadata</strong>: Additional page metadata</p>\n<h3 id=\"print_stats\">print_stats()</h3>\n<p>Display crawl statistics in formatted output.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">print_stats</span>(<span class=\"hljs-params\">\n    self,\n    detailed: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>\n</span>) -&gt; <span class=\"hljs-literal\">None</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"parameters_3\">Parameters</h4>\n<ul>\n<li><strong>detailed</strong> (<code>bool</code>): If True, shows detailed metrics with colors. If False, shows summary table.</li>\n</ul>\n<h3 id=\"export_knowledge_base\">export_knowledge_base()</h3>\n<p>Export the collected knowledge base to a JSONL file.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">export_knowledge_base</span>(<span class=\"hljs-params\">\n    self,\n    path: <span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, Path]\n</span>) -&gt; <span class=\"hljs-literal\">None</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"parameters_4\">Parameters</h4>\n<ul>\n<li><strong>path</strong> (<code>Union[str, Path]</code>): Output file path for JSONL export</li>\n</ul>\n<h4 id=\"example_1\">Example</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">adaptive.export_knowledge_base(<span class=\"hljs-string\">\"my_knowledge.jsonl\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"import_knowledge_base\">import_knowledge_base()</h3>\n<p>Import a previously exported knowledge base.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">import_knowledge_base</span>(<span class=\"hljs-params\">\n    self,\n    path: <span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, Path]\n</span>) -&gt; <span class=\"hljs-literal\">None</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"parameters_5\">Parameters</h4>\n<ul>\n<li><strong>path</strong> (<code>Union[str, Path]</code>): Path to JSONL file to import</li>\n</ul>\n<h2 id=\"configuration\">Configuration</h2>\n<p>The <code>AdaptiveConfig</code> class controls the behavior of adaptive crawling:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-meta\">@dataclass</span>\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">AdaptiveConfig</span>:\n    confidence_threshold: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.8</span>      <span class=\"hljs-comment\"># Stop when confidence reaches this</span>\n    max_pages: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">50</span>                    <span class=\"hljs-comment\"># Maximum pages to crawl</span>\n    top_k_links: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">5</span>                   <span class=\"hljs-comment\"># Links to follow per page</span>\n    min_gain_threshold: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.1</span>        <span class=\"hljs-comment\"># Minimum expected gain to continue</span>\n    save_state: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>               <span class=\"hljs-comment\"># Auto-save crawl state</span>\n    state_path: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>       <span class=\"hljs-comment\"># Path for state persistence</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"example-with-custom-config\">Example with Custom Config</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-built_in\">config</span> = AdaptiveConfig(\n    confidence_threshold=<span class=\"hljs-number\">0.7</span>,\n    max_pages=<span class=\"hljs-number\">20</span>,\n    top_k_links=<span class=\"hljs-number\">3</span>\n)\n\nadaptive = AdaptiveCrawler(crawler, <span class=\"hljs-built_in\">config</span>=<span class=\"hljs-built_in\">config</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"complete-example\">Complete Example</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, AdaptiveCrawler, AdaptiveConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Configure adaptive crawling</span>\n    config = AdaptiveConfig(\n        confidence_threshold=<span class=\"hljs-number\">0.75</span>,\n        max_pages=<span class=\"hljs-number\">15</span>,\n        save_state=<span class=\"hljs-literal\">True</span>,\n        state_path=<span class=\"hljs-string\">\"my_crawl.json\"</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        adaptive = AdaptiveCrawler(crawler, config)\n\n        <span class=\"hljs-comment\"># Start crawling</span>\n        state = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n            start_url=<span class=\"hljs-string\">\"https://example.com/docs\"</span>,\n            query=<span class=\"hljs-string\">\"authentication oauth2 jwt\"</span>\n        )\n\n        <span class=\"hljs-comment\"># Check results</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Confidence achieved: <span class=\"hljs-subst\">{adaptive.confidence:<span class=\"hljs-number\">.0</span>%}</span>\"</span>)\n        adaptive.print_stats()\n\n        <span class=\"hljs-comment\"># Get most relevant pages</span>\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> adaptive.get_relevant_content(top_k=<span class=\"hljs-number\">3</span>):\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"- <span class=\"hljs-subst\">{page[<span class=\"hljs-string\">'url'</span>]}</span> (score: <span class=\"hljs-subst\">{page[<span class=\"hljs-string\">'score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>)\"</span>)\n\n        <span class=\"hljs-comment\"># Export for later use</span>\n        adaptive.export_knowledge_base(<span class=\"hljs-string\">\"auth_knowledge.jsonl\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../digest/\">digest() Method Reference</a></li>\n<li><a href=\"../../core/adaptive-crawling/\">Adaptive Crawling Guide</a></li>\n<li><a href=\"../../advanced/adaptive-strategies/\">Advanced Adaptive Strategies</a></li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/arun/",
    "title": "arun() - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"arun-parameter-guide-new-approach\"><code>arun()</code> Parameter Guide (New Approach)</h1>\n<p>In Crawl4AI’s <strong>latest</strong> configuration model, nearly all parameters that once went directly to <code>arun()</code> are now part of <strong><code>CrawlerRunConfig</code></strong>. When calling <code>arun()</code>, you provide:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,  \n    config=my_run_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Below is an organized look at the parameters that can go inside <code>CrawlerRunConfig</code>, divided by their functional areas. For <strong>Browser</strong> settings (e.g., <code>headless</code>, <code>browser_type</code>), see <a href=\"../parameters/\">BrowserConfig</a>.</p>\n<hr>\n<h2 id=\"1-core-usage\">1. Core Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    run_config = CrawlerRunConfig(\n        verbose=<span class=\"hljs-literal\">True</span>,            <span class=\"hljs-comment\"># Detailed logging</span>\n        cache_mode=CacheMode.ENABLED,  <span class=\"hljs-comment\"># Use normal read/write cache</span>\n        check_robots_txt=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Respect robots.txt rules</span>\n        <span class=\"hljs-comment\"># ... other parameters</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n\n        <span class=\"hljs-comment\"># Check if blocked by robots.txt</span>\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success <span class=\"hljs-keyword\">and</span> result.status_code == <span class=\"hljs-number\">403</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Fields</strong>:\n- <code>verbose=True</code> logs each crawl step.  \n- <code>cache_mode</code> decides how to read/write the local crawl cache.</p>\n<hr>\n<h2 id=\"2-cache-control\">2. Cache Control</h2>\n<p><strong><code>cache_mode</code></strong> (default: <code>CacheMode.ENABLED</code>)<br>\nUse a built-in enum from <code>CacheMode</code>:</p>\n<ul>\n<li><code>ENABLED</code>: Normal caching—reads if available, writes if missing.</li>\n<li><code>DISABLED</code>: No caching—always refetch pages.</li>\n<li><code>READ_ONLY</code>: Reads from cache only; no new writes.</li>\n<li><code>WRITE_ONLY</code>: Writes to cache but doesn’t read existing data.</li>\n<li><code>BYPASS</code>: Skips reading cache for this crawl (though it might still write if set up that way).</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Cache modes</strong>:</p>\n<ul>\n<li><code>CacheMode.BYPASS</code> — Skip cache entirely; always fetch fresh, write result to cache.</li>\n<li><code>CacheMode.DISABLED</code> — No caching at all; don't read or write.</li>\n<li><code>CacheMode.WRITE_ONLY</code> — Never read from cache, but write results.</li>\n<li><code>CacheMode.READ_ONLY</code> — Read from cache if available, never write.</li>\n</ul>\n<hr>\n<h2 id=\"3-content-processing-selection\">3. Content Processing &amp; Selection</h2>\n<h3 id=\"31-text-processing\">3.1 Text Processing</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">10</span>,   <span class=\"hljs-comment\"># Ignore text blocks &lt;10 words</span>\n    only_text<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>,           <span class=\"hljs-comment\"># If True, tries to remove non-text elements</span>\n    keep_data_attributes<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span> <span class=\"hljs-comment\"># Keep or discard data-* attributes</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-content-selection\">3.2 Content Selection</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    css_selector<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\".main-content\"</span>,  <span class=\"hljs-comment\"># Focus on .main-content region only</span>\n    excluded_tags<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"form\"</span>, <span class=\"hljs-string\">\"nav\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Remove entire tag blocks</span>\n    remove_forms<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># Specifically strip &lt;form&gt; elements</span>\n    remove_overlay_elements<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Attempt to remove modals/popups</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"33-link-handling\">3.3 Link Handling</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,         <span class=\"hljs-comment\"># Remove external links from final content</span>\n    exclude_social_media_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,     <span class=\"hljs-comment\"># Remove links to known social sites</span>\n    exclude_domains<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"ads.example.com\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Exclude links to these domains</span>\n    exclude_social_media_domains<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"facebook.com\"</span>,<span class=\"hljs-string\">\"twitter.com\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Extend the default list</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"34-media-filtering\">3.4 Media Filtering</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Strip images from other domains</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"4-page-navigation-timing\">4. Page Navigation &amp; Timing</h2>\n<h3 id=\"41-basic-browser-flow\">4.1 Basic Browser Flow</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    wait_for=<span class=\"hljs-string\">\"css:.dynamic-content\"</span>, <span class=\"hljs-comment\"># Wait for .dynamic-content</span>\n    delay_before_return_html=2.0,    <span class=\"hljs-comment\"># Wait 2s before capturing final HTML</span>\n    page_timeout=60000,             <span class=\"hljs-comment\"># Navigation &amp; script timeout (ms)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Fields</strong>:</p>\n<ul>\n<li><code>wait_for</code>:  </li>\n<li><code>\"css:selector\"</code> or  </li>\n<li>\n<p><code>\"js:() =&gt; boolean\"</code><br>\n  e.g. <code>js:() =&gt; document.querySelectorAll('.item').length &gt; 10</code>.</p>\n</li>\n<li>\n<p><code>mean_delay</code> &amp; <code>max_range</code>: define random delays for <code>arun_many()</code> calls.  </p>\n</li>\n<li><code>semaphore_count</code>: concurrency limit when crawling multiple URLs.</li>\n</ul>\n<h3 id=\"42-javascript-execution\">4.2 JavaScript Execution</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    js_code<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        <span class=\"hljs-string\">\"document.querySelector('.load-more')?.click();\"</span>\n    <span class=\"hljs-punctuation\">]</span>,\n    js_only<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ul>\n<li><code>js_code</code> can be a single string or a list of strings.  </li>\n<li><code>js_only=True</code> means “I’m continuing in the same session with new JS steps, no new full navigation.”</li>\n</ul>\n<h3 id=\"43-anti-bot\">4.3 Anti-Bot</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    magic<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    simulate_user<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    override_navigator<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>magic=True</code> tries multiple stealth features.  \n- <code>simulate_user=True</code> mimics mouse movements or random delays.  \n- <code>override_navigator=True</code> fakes some navigator properties (like user agent checks).<p></p>\n<hr>\n<h2 id=\"5-session-management\">5. Session Management</h2>\n<p><strong><code>session_id</code></strong>: \n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    session_id=<span class=\"hljs-string\">\"my_session123\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nIf re-used in subsequent <code>arun()</code> calls, the same tab/page context is continued (helpful for multi-step tasks or stateful browsing).<p></p>\n<hr>\n<h2 id=\"6-screenshot-pdf-media-options\">6. Screenshot, PDF &amp; Media Options</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    screenshot<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># Grab a screenshot as base64</span>\n    screenshot_wait_for<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">1.0</span>,     <span class=\"hljs-comment\"># Wait 1s before capturing</span>\n    pdf<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Also produce a PDF</span>\n    image_description_min_word_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5</span>,  <span class=\"hljs-comment\"># If analyzing alt text</span>\n    image_score_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3</span>,                <span class=\"hljs-comment\"># Filter out low-score images</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Where they appear</strong>:\n- <code>result.screenshot</code> → Base64 screenshot string.\n- <code>result.pdf</code> → Byte array with PDF data.<p></p>\n<hr>\n<h2 id=\"7-extraction-strategy\">7. Extraction Strategy</h2>\n<p><strong>For advanced data extraction</strong> (CSS/LLM-based), set <code>extraction_strategy</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    extraction_strategy=my_css_or_llm_strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The extracted data will appear in <code>result.extracted_content</code>.</p>\n<hr>\n<h2 id=\"8-comprehensive-example\">8. Comprehensive Example</h2>\n<p>Below is a snippet combining many parameters:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Example schema</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"article.post\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>,  <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>,  <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n\n    run_config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Core</span>\n        verbose=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.ENABLED,\n        check_robots_txt=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Respect robots.txt rules</span>\n\n        <span class=\"hljs-comment\"># Content</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        css_selector=<span class=\"hljs-string\">\"main.content\"</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Page &amp; JS</span>\n        js_code=<span class=\"hljs-string\">\"document.querySelector('.show-more')?.click();\"</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.loaded-block\"</span>,\n        page_timeout=<span class=\"hljs-number\">30000</span>,\n\n        <span class=\"hljs-comment\"># Extraction</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n\n        <span class=\"hljs-comment\"># Session</span>\n        session_id=<span class=\"hljs-string\">\"persistent_session\"</span>,\n\n        <span class=\"hljs-comment\"># Media</span>\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        pdf=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Anti-bot</span>\n        simulate_user=<span class=\"hljs-literal\">True</span>,\n        magic=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/posts\"</span>, config=run_config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extraction JSON:\"</span>, result.extracted_content)\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n            <span class=\"hljs-keyword\">if</span> result.pdf:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"PDF bytes length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.pdf))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What we covered</strong>:</p>\n<p>1. <strong>Crawling</strong> the main content region, ignoring external links.  \n2. Running <strong>JavaScript</strong> to click “.show-more”.  \n3. <strong>Waiting</strong> for “.loaded-block” to appear.  \n4. Generating a <strong>screenshot</strong> &amp; <strong>PDF</strong> of the final page.  \n5. Extracting repeated “article.post” elements with a <strong>CSS-based</strong> extraction strategy.</p>\n<hr>\n<h2 id=\"9-best-practices\">9. Best Practices</h2>\n<p>1. <strong>Use <code>BrowserConfig</code> for global browser</strong> settings (headless, user agent).  \n2. <strong>Use <code>CrawlerRunConfig</code></strong> to handle the <strong>specific</strong> crawl needs: content filtering, caching, JS, screenshot, extraction, etc.  \n3. Keep your <strong>parameters consistent</strong> in run configs—especially if you’re part of a large codebase with multiple crawls.  \n4. <strong>Limit</strong> large concurrency (<code>semaphore_count</code>) if the site or your system can’t handle it.  \n5. For dynamic pages, set <code>js_code</code> or <code>scan_full_page</code> so you load all content.</p>\n<hr>\n<h2 id=\"10-conclusion\">10. Conclusion</h2>\n<p>All parameters that used to be direct arguments to <code>arun()</code> now belong in <strong><code>CrawlerRunConfig</code></strong>. This approach:</p>\n<ul>\n<li>Makes code <strong>clearer</strong> and <strong>more maintainable</strong>.  </li>\n<li>Minimizes confusion about which arguments affect global vs. per-crawl behavior.  </li>\n<li>Allows you to create <strong>reusable</strong> config objects for different pages or tasks.</li>\n</ul>\n<p>For a <strong>full</strong> reference, check out the <a href=\"../parameters/\">CrawlerRunConfig Docs</a>. </p>\n<p>Happy crawling with your <strong>structured, flexible</strong> config approach!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/arun_many/",
    "title": "arun_many() - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"arun_many-reference\"><code>arun_many(...)</code> Reference</h1>\n<blockquote>\n<p><strong>Note</strong>: This function is very similar to <a href=\"../arun/\"><code>arun()</code></a> but focused on <strong>concurrent</strong> or <strong>batch</strong> crawling. If you’re unfamiliar with <code>arun()</code> usage, please read that doc first, then review this for differences.</p>\n</blockquote>\n<h2 id=\"function-signature\">Function Signature</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun_many</span>(<span class=\"hljs-params\">\n    urls: <span class=\"hljs-type\">Union</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>], <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Any</span>]],\n    config: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[CrawlerRunConfig, <span class=\"hljs-type\">List</span>[CrawlerRunConfig]]] = <span class=\"hljs-literal\">None</span>,\n    dispatcher: <span class=\"hljs-type\">Optional</span>[BaseDispatcher] = <span class=\"hljs-literal\">None</span>,\n    ...\n</span>) -&gt; RunManyReturn:\n    <span class=\"hljs-string\">\"\"\"\n    Crawl multiple URLs concurrently or in batches.\n\n    :param urls: A list of URLs (or tasks) to crawl.\n    :param config: (Optional) Either:\n        - A single `CrawlerRunConfig` applying to all URLs\n        - A list of `CrawlerRunConfig` objects with url_matcher patterns\n    :param dispatcher: (Optional) A concurrency controller (e.g. MemoryAdaptiveDispatcher).\n    ...\n    :return: RunManyReturn containing either a list of `CrawlResult` objects or an async generator if streaming is enabled.\n    \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"differences-from-arun\">Differences from <code>arun()</code></h2>\n<p>1. <strong>Multiple URLs</strong>:  </p>\n<ul>\n<li>Instead of crawling a single URL, you pass a list of them (strings or tasks).  </li>\n<li>The function returns <code>RunManyReturn</code> which contains either a <strong>list</strong> of <code>CrawlResult</code> or an <strong>async generator</strong> if streaming is enabled.</li>\n</ul>\n<p>2. <strong>Concurrency &amp; Dispatchers</strong>:  </p>\n<ul>\n<li><strong><code>dispatcher</code></strong> param allows advanced concurrency control.  </li>\n<li>If omitted, a default dispatcher (like <code>MemoryAdaptiveDispatcher</code>) is used internally.  </li>\n<li>Dispatchers handle concurrency, rate limiting, and memory-based adaptive throttling (see <a href=\"../../advanced/multi-url-crawling/\">Multi-URL Crawling</a>).</li>\n</ul>\n<p>3. <strong>Streaming Support</strong>:  </p>\n<ul>\n<li>Enable streaming by setting <code>stream=True</code> in your <code>CrawlerRunConfig</code>.</li>\n<li>When streaming, use <code>async for</code> to process results as they become available.</li>\n<li>Ideal for processing large numbers of URLs without waiting for all to complete.</li>\n</ul>\n<p>4. <strong>Parallel</strong> Execution**:  </p>\n<ul>\n<li><code>arun_many()</code> can run multiple requests concurrently under the hood.  </li>\n<li>Each <code>CrawlResult</code> might also include a <strong><code>dispatch_result</code></strong> with concurrency details (like memory usage, start/end times).</li>\n</ul>\n<h3 id=\"basic-example-batch-mode\">Basic Example (Batch Mode)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Minimal usage: The default dispatcher will be used</span>\nresults = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>],\n    config=CrawlerRunConfig(stream=<span class=\"hljs-literal\">False</span>)  <span class=\"hljs-comment\"># Default behavior</span>\n)\n\n<span class=\"hljs-keyword\">for</span> res <span class=\"hljs-keyword\">in</span> results:\n    <span class=\"hljs-keyword\">if</span> res.success:\n        <span class=\"hljs-built_in\">print</span>(res.url, <span class=\"hljs-string\">\"crawled OK!\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Failed:\"</span>, res.url, <span class=\"hljs-string\">\"-\"</span>, res.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"streaming-example\">Streaming Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    stream=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    cache_mode=CacheMode.BYPASS\n)\n\n<span class=\"hljs-comment\"># Process results as they complete</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>, <span class=\"hljs-string\">\"https://site3.com\"</span>],\n    config=config\n):\n    <span class=\"hljs-keyword\">if</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Just completed: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n        <span class=\"hljs-comment\"># Process each result immediately</span>\n        process_result(result)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"with-a-custom-dispatcher\">With a Custom Dispatcher</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">dispatcher = MemoryAdaptiveDispatcher(\n    memory_threshold_percent=70.0,\n    max_session_permit=10\n)\nresults = await crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>, <span class=\"hljs-string\">\"https://site3.com\"</span>],\n    config=my_run_config,\n    dispatcher=dispatcher\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"url-specific-configurations\">URL-Specific Configurations</h3>\n<p>Instead of using one config for all URLs, provide a list of configs with <code>url_matcher</code> patterns:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, MatchMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.extraction_strategy <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-comment\"># PDF files - specialized extraction</span>\npdf_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>,\n    scraping_strategy=PDFContentScrapingStrategy()\n)\n\n<span class=\"hljs-comment\"># Blog/article pages - content filtering</span>\nblog_config = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"*/blog/*\"</span>, <span class=\"hljs-string\">\"*/article/*\"</span>, <span class=\"hljs-string\">\"*python.org*\"</span>],\n    markdown_generator=DefaultMarkdownGenerator(\n        content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.48</span>)\n    )\n)\n\n<span class=\"hljs-comment\"># Dynamic pages - JavaScript execution</span>\ngithub_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'github.com'</span> <span class=\"hljs-keyword\">in</span> url,\n    js_code=<span class=\"hljs-string\">\"window.scrollTo(0, 500);\"</span>\n)\n\n<span class=\"hljs-comment\"># API endpoints - JSON extraction</span>\napi_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'api'</span> <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">or</span> url.endswith(<span class=\"hljs-string\">'.json'</span>),\n    <span class=\"hljs-comment\"># Custome settings for JSON extraction</span>\n)\n\n<span class=\"hljs-comment\"># Default fallback config</span>\ndefault_config = CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher means it never matches except as fallback</span>\n\n<span class=\"hljs-comment\"># Pass the list of configs - first match wins!</span>\nresults = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[\n        <span class=\"hljs-string\">\"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf\"</span>,  <span class=\"hljs-comment\"># → pdf_config</span>\n        <span class=\"hljs-string\">\"https://blog.python.org/\"</span>,  <span class=\"hljs-comment\"># → blog_config</span>\n        <span class=\"hljs-string\">\"https://github.com/microsoft/playwright\"</span>,  <span class=\"hljs-comment\"># → github_config</span>\n        <span class=\"hljs-string\">\"https://httpbin.org/json\"</span>,  <span class=\"hljs-comment\"># → api_config</span>\n        <span class=\"hljs-string\">\"https://example.com/\"</span>  <span class=\"hljs-comment\"># → default_config</span>\n    ],\n    config=[pdf_config, blog_config, github_config, api_config, default_config]\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>URL Matching Features</strong>:\n- <strong>String patterns</strong>: <code>\"*.pdf\"</code>, <code>\"*/blog/*\"</code>, <code>\"*python.org*\"</code>\n- <strong>Function matchers</strong>: <code>lambda url: 'api' in url</code>\n- <strong>Mixed patterns</strong>: Combine strings and functions with <code>MatchMode.OR</code> or <code>MatchMode.AND</code>\n- <strong>First match wins</strong>: Configs are evaluated in order</p>\n<p><strong>Key Points</strong>:\n- Each URL is processed by the same or separate sessions, depending on the dispatcher’s strategy.\n- <code>dispatch_result</code> in each <code>CrawlResult</code> (if using concurrency) can hold memory and timing info.  \n- If you need to handle authentication or session IDs, pass them in each individual task or within your run config.\n- <strong>Important</strong>: Always include a default config (without <code>url_matcher</code>) as the last item if you want to handle all URLs. Otherwise, unmatched URLs will fail.</p>\n<h3 id=\"return-value\">Return Value</h3>\n<p>Returns a <strong><code>RunManyReturn</code></strong> object which contains either a <strong>list</strong> of <a href=\"../crawl-result/\"><code>CrawlResult</code></a> objects, or an <strong>async generator</strong> if streaming is enabled. You can iterate to check <code>result.success</code> or read each item’s <code>extracted_content</code>, <code>markdown</code>, or <code>dispatch_result</code>.</p>\n<hr>\n<h2 id=\"dispatcher-reference\">Dispatcher Reference</h2>\n<ul>\n<li><strong><code>MemoryAdaptiveDispatcher</code></strong>: Dynamically manages concurrency based on system memory usage.  </li>\n<li><strong><code>SemaphoreDispatcher</code></strong>: Fixed concurrency limit, simpler but less adaptive.  </li>\n</ul>\n<p>For advanced usage or custom settings, see <a href=\"../../advanced/multi-url-crawling/\">Multi-URL Crawling with Dispatchers</a>.</p>\n<hr>\n<h2 id=\"common-pitfalls\">Common Pitfalls</h2>\n<p>1. <strong>Large Lists</strong>: If you pass thousands of URLs, be mindful of memory or rate-limits. A dispatcher can help.  </p>\n<p>2. <strong>Session Reuse</strong>: If you need specialized logins or persistent contexts, ensure your dispatcher or tasks handle sessions accordingly.  </p>\n<p>3. <strong>Error Handling</strong>: Each <code>CrawlResult</code> might fail for different reasons—always check <code>result.success</code> or the <code>error_message</code> before proceeding.</p>\n<hr>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Use <code>arun_many()</code> when you want to <strong>crawl multiple URLs</strong> simultaneously or in controlled parallel tasks. If you need advanced concurrency features (like memory-based adaptive throttling or complex rate-limiting), provide a <strong>dispatcher</strong>. Each result is a standard <code>CrawlResult</code>, possibly augmented with concurrency stats (<code>dispatch_result</code>) for deeper inspection. For more details on concurrency logic and dispatchers, see the <a href=\"../../advanced/multi-url-crawling/\">Advanced Multi-URL Crawling</a> docs.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/async-webcrawler/",
    "title": "AsyncWebCrawler - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"asyncwebcrawler\">AsyncWebCrawler</h1>\n<p>The <strong><code>AsyncWebCrawler</code></strong> is the core class for asynchronous web crawling in Crawl4AI. You typically create it <strong>once</strong>, optionally customize it with a <strong><code>BrowserConfig</code></strong> (e.g., headless, user agent), then <strong>run</strong> multiple <strong><code>arun()</code></strong> calls with different <strong><code>CrawlerRunConfig</code></strong> objects.</p>\n<p><strong>Recommended usage</strong>:</p>\n<p>1. <strong>Create</strong> a <code>BrowserConfig</code> for global browser settings.  </p>\n<p>2. <strong>Instantiate</strong> <code>AsyncWebCrawler(config=browser_config)</code>.  </p>\n<p>3. <strong>Use</strong> the crawler in an async context manager (<code>async with</code>) or manage start/close manually.  </p>\n<p>4. <strong>Call</strong> <code>arun(url, config=crawler_run_config)</code> for each page you want.</p>\n<hr>\n<h2 id=\"1-constructor-overview\">1. Constructor Overview</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">AsyncWebCrawler</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        self,\n        crawler_strategy: <span class=\"hljs-type\">Optional</span>[AsyncCrawlerStrategy] = <span class=\"hljs-literal\">None</span>,\n        config: <span class=\"hljs-type\">Optional</span>[BrowserConfig] = <span class=\"hljs-literal\">None</span>,\n        always_bypass_cache: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>,           <span class=\"hljs-comment\"># deprecated</span>\n        always_by_pass_cache: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">bool</span>] = <span class=\"hljs-literal\">None</span>, <span class=\"hljs-comment\"># also deprecated</span>\n        base_directory: <span class=\"hljs-built_in\">str</span> = ...,\n        thread_safe: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>,\n        **kwargs,\n    </span>):\n        <span class=\"hljs-string\">\"\"\"\n        Create an AsyncWebCrawler instance.\n\n        Args:\n            crawler_strategy: \n                (Advanced) Provide a custom crawler strategy if needed.\n            config: \n                A BrowserConfig object specifying how the browser is set up.\n            always_bypass_cache: \n                (Deprecated) Use CrawlerRunConfig.cache_mode instead.\n            base_directory:     \n                Folder for storing caches/logs (if relevant).\n            thread_safe: \n                If True, attempts some concurrency safeguards. Usually False.\n            **kwargs: \n                Additional legacy or debugging parameters.\n        \"\"\"</span>\n    )\n\n<span class=\"hljs-comment\">### Typical Initialization</span>\n\n```python\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nbrowser_cfg = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n    headless=<span class=\"hljs-literal\">True</span>,\n    verbose=<span class=\"hljs-literal\">True</span>\n)\n\ncrawler = AsyncWebCrawler(config=browser_cfg)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Notes</strong>:</p>\n<ul>\n<li><strong>Legacy</strong> parameters like <code>always_bypass_cache</code> remain for backward compatibility, but prefer to set <strong>caching</strong> in <code>CrawlerRunConfig</code>.</li>\n</ul>\n<hr>\n<h2 id=\"2-lifecycle-startclose-or-context-manager\">2. Lifecycle: Start/Close or Context Manager</h2>\n<h3 id=\"21-context-manager-recommended\">2.1 Context Manager (Recommended)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>(<span class=\"hljs-params\">config=browser_cfg</span>) <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n    <span class=\"hljs-meta\"># The crawler automatically starts/closes resources</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>When the <code>async with</code> block ends, the crawler cleans up (closes the browser, etc.).</p>\n<h3 id=\"22-manual-start-close\">2.2 Manual Start &amp; Close</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">crawler = AsyncWebCrawler(config=browser_cfg)\n<span class=\"hljs-keyword\">await</span> crawler.start()\n\nresult1 = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\nresult2 = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://another.com\"</span>)\n\n<span class=\"hljs-keyword\">await</span> crawler.close()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Use this style if you have a <strong>long-running</strong> application or need full control of the crawler’s lifecycle.</p>\n<hr>\n<h2 id=\"3-primary-method-arun\">3. Primary Method: <code>arun()</code></h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun</span>(<span class=\"hljs-params\">\n    self,\n    url: <span class=\"hljs-built_in\">str</span>,\n    config: <span class=\"hljs-type\">Optional</span>[CrawlerRunConfig] = <span class=\"hljs-literal\">None</span>,\n    <span class=\"hljs-comment\"># Legacy parameters for backward compatibility...</span>\n</span>) -&gt; RunManyReturn:\n    ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"31-new-approach\">3.1 New Approach</h3>\n<p>You pass a <code>CrawlerRunConfig</code> object that sets up everything about a crawl—content filtering, caching, session reuse, JS code, screenshots, etc.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, CacheMode\n\nrun_cfg = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS,\n    css_selector=<span class=\"hljs-string\">\"main.article\"</span>,\n    word_count_threshold=<span class=\"hljs-number\">10</span>,\n    screenshot=<span class=\"hljs-literal\">True</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/news\"</span>, config=run_cfg)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawled HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n    <span class=\"hljs-keyword\">if</span> result.screenshot:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot base64 length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-legacy-parameters-still-accepted\">3.2 Legacy Parameters Still Accepted</h3>\n<p>For <strong>backward</strong> compatibility, <code>arun()</code> can still accept direct arguments like <code>css_selector=...</code>, <code>word_count_threshold=...</code>, etc., but we strongly advise migrating them into a <strong><code>CrawlerRunConfig</code></strong>.</p>\n<hr>\n<h2 id=\"4-batch-processing-arun_many\">4. Batch Processing: <code>arun_many()</code></h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun_many</span>(<span class=\"hljs-params\">\n    self,\n    urls: <span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>],\n    config: <span class=\"hljs-type\">Optional</span>[CrawlerRunConfig] = <span class=\"hljs-literal\">None</span>,\n    <span class=\"hljs-comment\"># Legacy parameters maintained for backwards compatibility...</span>\n</span>) -&gt; RunManyReturn:\n    <span class=\"hljs-string\">\"\"\"\n    Process multiple URLs with intelligent rate limiting and resource monitoring.\n    \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"41-resource-aware-crawling\">4.1 Resource-Aware Crawling</h3>\n<p>The <code>arun_many()</code> method now uses an intelligent dispatcher that:</p>\n<ul>\n<li>Monitors system memory usage</li>\n<li>Implements adaptive rate limiting</li>\n<li>Provides detailed progress monitoring</li>\n<li>Manages concurrent crawls efficiently</li>\n</ul>\n<h3 id=\"42-example-usage\">4.2 Example Usage</h3>\n<p>Check page <a href=\"../../advanced/multi-url-crawling/\">Multi-url Crawling</a> for a detailed example of how to use <code>arun_many()</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\">### 4.3 Key Features</span>\n\n<span class=\"hljs-number\">1.</span> **Rate Limiting**\n\n   - Automatic delay between requests\n   - Exponential backoff on rate limit detection\n   - Domain-specific rate limiting\n   - Configurable retry strategy\n\n<span class=\"hljs-number\">2.</span> **Resource Monitoring**\n\n   - Memory usage tracking\n   - Adaptive concurrency based on system load\n   - Automatic pausing when resources are constrained\n\n<span class=\"hljs-number\">3.</span> **Progress Monitoring**\n\n   - Detailed <span class=\"hljs-keyword\">or</span> aggregated progress display\n   - Real-time status updates\n   - Memory usage statistics\n\n<span class=\"hljs-number\">4.</span> **Error Handling**\n\n   - Graceful handling of rate limits\n   - Automatic retries <span class=\"hljs-keyword\">with</span> backoff\n   - Detailed error reporting\n\n---\n\n<span class=\"hljs-comment\">## 5. `CrawlResult` Output</span>\n\nEach `arun()` returns a **`CrawlResult`** containing:\n\n- `url`: Final URL (<span class=\"hljs-keyword\">if</span> redirected).\n- `html`: Original HTML.\n- `cleaned_html`: Sanitized HTML.\n- `markdown_v2`: Removed <span class=\"hljs-keyword\">in</span> v0<span class=\"hljs-number\">.5</span>. Accessing it raises `AttributeError`; use `markdown`.\n- `extracted_content`: If an extraction strategy was used (JSON <span class=\"hljs-keyword\">for</span> CSS/LLM strategies).\n- `screenshot`, `pdf`: If screenshots/PDF requested.\n- `media`, `links`: Information about discovered images/links.\n- `success`, `error_message`: Status info.\n\nFor details, see [CrawlResult doc](./crawl-result.md).\n\n---\n\n<span class=\"hljs-comment\">## 6. Quick Example</span>\n\nBelow <span class=\"hljs-keyword\">is</span> an example hooking it <span class=\"hljs-built_in\">all</span> together:\n\n```python\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">import</span> json\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Browser config</span>\n    browser_cfg = BrowserConfig(\n        browser_type=<span class=\"hljs-string\">\"firefox\"</span>,\n        headless=<span class=\"hljs-literal\">False</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># 2. Run config</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"article.post\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"url\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, \n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>\n            }\n        ]\n    }\n\n    run_cfg = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n        word_count_threshold=<span class=\"hljs-number\">15</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.post\"</span>  <span class=\"hljs-comment\"># Wait for posts to appear</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/blog\"</span>,\n            config=run_cfg\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-keyword\">if</span> result.extracted_content:\n                articles = json.loads(result.extracted_content)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted articles:\"</span>, articles[:<span class=\"hljs-number\">2</span>])\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Explanation</strong>:</p>\n<ul>\n<li>We define a <strong><code>BrowserConfig</code></strong> with Firefox, no headless, and <code>verbose=True</code>.  </li>\n<li>We define a <strong><code>CrawlerRunConfig</code></strong> that <strong>bypasses cache</strong>, uses a <strong>CSS</strong> extraction schema, has a <code>word_count_threshold=15</code>, etc.  </li>\n<li>We pass them to <code>AsyncWebCrawler(config=...)</code> and <code>arun(url=..., config=...)</code>.</li>\n</ul>\n<hr>\n<h2 id=\"7-best-practices-migration-notes\">7. Best Practices &amp; Migration Notes</h2>\n<p>1. <strong>Use</strong> <code>BrowserConfig</code> for <strong>global</strong> settings about the browser’s environment.  \n2. <strong>Use</strong> <code>CrawlerRunConfig</code> for <strong>per-crawl</strong> logic (caching, content filtering, extraction strategies, wait conditions).  \n3. <strong>Avoid</strong> legacy parameters like <code>css_selector</code> or <code>word_count_threshold</code> directly in <code>arun()</code>. Instead:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">run_cfg</span> = CrawlerRunConfig(css_selector=<span class=\"hljs-string\">\".main-content\"</span>, word_count_threshold=<span class=\"hljs-number\">20</span>)\n<span class=\"hljs-attr\">result</span> = await crawler.arun(url=<span class=\"hljs-string\">\"...\"</span>, config=run_cfg)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>4. <strong>Context Manager</strong> usage is simplest unless you want a persistent crawler across many calls.</p>\n<hr>\n<h2 id=\"8-summary\">8. Summary</h2>\n<p><strong>AsyncWebCrawler</strong> is your entry point to asynchronous crawling:</p>\n<ul>\n<li><strong>Constructor</strong> accepts <strong><code>BrowserConfig</code></strong> (or defaults).  </li>\n<li><strong><code>arun(url, config=CrawlerRunConfig)</code></strong> is the main method for single-page crawls.  </li>\n<li><strong><code>arun_many(urls, config=CrawlerRunConfig)</code></strong> handles concurrency across multiple URLs.  </li>\n<li>For advanced lifecycle control, use <code>start()</code> and <code>close()</code> explicitly.  </li>\n</ul>\n<p><strong>Migration</strong>:  </p>\n<ul>\n<li>If you used <code>AsyncWebCrawler(browser_type=\"chromium\", css_selector=\"...\")</code>, move browser settings to <code>BrowserConfig(...)</code> and content/crawl logic to <code>CrawlerRunConfig(...)</code>.</li>\n</ul>\n<p>This modular approach ensures your code is <strong>clean</strong>, <strong>scalable</strong>, and <strong>easy to maintain</strong>. For any advanced or rarely used parameters, see the <a href=\"../parameters/\">BrowserConfig docs</a>.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/c4a-script-reference/",
    "title": "C4A-Script Reference - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"c4a-script-api-reference\">C4A-Script API Reference</h1>\n<p>Complete reference for all C4A-Script commands, syntax, and advanced features.</p>\n<h2 id=\"command-categories\">Command Categories</h2>\n<h3 id=\"navigation-commands\">🧭 Navigation Commands</h3>\n<p>Navigate between pages and manage browser history.</p>\n<h4 id=\"go-url\"><code>GO &lt;url&gt;</code></h4>\n<p>Navigate to a specific URL.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">GO <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">url</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>url</code> - Target URL (string)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">GO https://example.com\nGO https://api.example.com/login\nGO /relative/path\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Supports both absolute and relative URLs\n- Automatically handles protocol detection\n- Waits for page load to complete</p>\n<hr>\n<h4 id=\"reload\"><code>RELOAD</code></h4>\n<p>Refresh the current page.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">RELOAD\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">RELOAD\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Equivalent to pressing F5 or clicking browser refresh\n- Waits for page reload to complete\n- Preserves current URL</p>\n<hr>\n<h4 id=\"back\"><code>BACK</code></h4>\n<p>Navigate back in browser history.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">BACK\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">BACK\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Equivalent to clicking browser back button\n- Does nothing if no previous page exists\n- Waits for navigation to complete</p>\n<hr>\n<h4 id=\"forward\"><code>FORWARD</code></h4>\n<p>Navigate forward in browser history.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">FORWARD\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">FORWARD\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Equivalent to clicking browser forward button\n- Does nothing if no next page exists\n- Waits for navigation to complete</p>\n<h3 id=\"wait-commands\">⏱️ Wait Commands</h3>\n<p>Control timing and synchronization with page elements.</p>\n<h4 id=\"wait-time\"><code>WAIT &lt;time&gt;</code></h4>\n<p>Wait for a specified number of seconds.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">WAIT <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">seconds</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>seconds</code> - Number of seconds to wait (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">WAIT <span class=\"hljs-number\">3</span>\nWAIT <span class=\"hljs-number\">1.5</span>\nWAIT <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Accepts decimal values\n- Useful for giving dynamic content time to load\n- Non-blocking for other browser operations</p>\n<hr>\n<h4 id=\"wait-selector-timeout\"><code>WAIT &lt;selector&gt; &lt;timeout&gt;</code></h4>\n<p>Wait for an element to appear on the page.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">WAIT `<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">selector</span>&gt;</span>` <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">timeout</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for the element (string in backticks)\n- <code>timeout</code> - Maximum seconds to wait (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">WAIT `<span class=\"hljs-selector-id\">#content</span>` <span class=\"hljs-number\">10</span>\nWAIT `<span class=\"hljs-selector-class\">.loading-spinner</span>` <span class=\"hljs-number\">5</span>\nWAIT `<span class=\"hljs-selector-tag\">button</span><span class=\"hljs-selector-attr\">[type=<span class=\"hljs-string\">\"submit\"</span>]</span>` <span class=\"hljs-number\">15</span>\nWAIT `<span class=\"hljs-selector-class\">.results</span> <span class=\"hljs-selector-class\">.item</span><span class=\"hljs-selector-pseudo\">:first</span>-child` <span class=\"hljs-number\">8</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Fails if element doesn't appear within timeout\n- More reliable than fixed time waits\n- Supports complex CSS selectors</p>\n<hr>\n<h4 id=\"wait-text-timeout\"><code>WAIT \"&lt;text&gt;\" &lt;timeout&gt;</code></h4>\n<p>Wait for specific text to appear anywhere on the page.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">WAIT <span class=\"hljs-string\">\"&lt;text&gt;\"</span> &lt;<span class=\"hljs-built_in\">timeout</span>&gt;\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>text</code> - Text content to wait for (string in quotes)\n- <code>timeout</code> - Maximum seconds to wait (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">WAIT <span class=\"hljs-string\">\"Loading complete\"</span> 10\nWAIT <span class=\"hljs-string\">\"Welcome back\"</span> 5\nWAIT <span class=\"hljs-string\">\"Search results\"</span> 15\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Case-sensitive text matching\n- Searches entire page content\n- Useful for dynamic status messages</p>\n<h3 id=\"mouse-commands\">🖱️ Mouse Commands</h3>\n<p>Simulate mouse interactions and movements.</p>\n<h4 id=\"click-selector\"><code>CLICK &lt;selector&gt;</code></h4>\n<p>Click on an element specified by CSS selector.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">CLICK <span class=\"hljs-string\">`&lt;selector&gt;`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for the element (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-less\"><span class=\"hljs-selector-tag\">CLICK</span> `<span class=\"hljs-selector-id\">#submit-button</span>`\n<span class=\"hljs-selector-tag\">CLICK</span> `<span class=\"hljs-selector-class\">.menu-item</span><span class=\"hljs-selector-pseudo\">:first</span><span class=\"hljs-selector-tag\">-child</span>`\n<span class=\"hljs-selector-tag\">CLICK</span> `<span class=\"hljs-selector-tag\">button</span><span class=\"hljs-selector-attr\">[data-action=<span class=\"hljs-string\">\"save\"</span>]</span>`\n<span class=\"hljs-selector-tag\">CLICK</span> `<span class=\"hljs-selector-tag\">a</span><span class=\"hljs-selector-attr\">[href=<span class=\"hljs-string\">\"/dashboard\"</span>]</span>`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Waits for element to be clickable\n- Scrolls element into view if necessary\n- Handles overlapping elements intelligently</p>\n<hr>\n<h4 id=\"click-x-y\"><code>CLICK &lt;x&gt; &lt;y&gt;</code></h4>\n<p>Click at specific coordinates on the page.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">CLICK <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">x</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">y</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>x</code> - X coordinate in pixels (number)\n- <code>y</code> - Y coordinate in pixels (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-objectivec\"><span class=\"hljs-built_in\">CLICK</span> <span class=\"hljs-number\">100</span> <span class=\"hljs-number\">200</span>\n<span class=\"hljs-built_in\">CLICK</span> <span class=\"hljs-number\">500</span> <span class=\"hljs-number\">300</span>\n<span class=\"hljs-built_in\">CLICK</span> <span class=\"hljs-number\">0</span> <span class=\"hljs-number\">0</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Coordinates are relative to viewport\n- Useful when element selectors are unreliable\n- Consider responsive design implications</p>\n<hr>\n<h4 id=\"double_click-selector\"><code>DOUBLE_CLICK &lt;selector&gt;</code></h4>\n<p>Double-click on an element.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">DOUBLE_CLICK <span class=\"hljs-string\">`&lt;selector&gt;`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for the element (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">DOUBLE_CLICK <span class=\"hljs-string\">`.file-icon`</span>\nDOUBLE_CLICK <span class=\"hljs-string\">`#editable-cell`</span>\nDOUBLE_CLICK <span class=\"hljs-string\">`.expandable-item`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Triggers dblclick event\n- Common for opening files or editing inline content\n- Timing between clicks is automatically handled</p>\n<hr>\n<h4 id=\"right_click-selector\"><code>RIGHT_CLICK &lt;selector&gt;</code></h4>\n<p>Right-click on an element to open context menu.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">RIGHT_CLICK <span class=\"hljs-string\">`&lt;selector&gt;`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for the element (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">RIGHT_CLICK <span class=\"hljs-string\">`#context-target`</span>\nRIGHT_CLICK <span class=\"hljs-string\">`.menu-trigger`</span>\nRIGHT_CLICK <span class=\"hljs-string\">`img.thumbnail`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Opens browser/application context menu\n- Useful for testing context menu interactions\n- May be blocked by some applications</p>\n<hr>\n<h4 id=\"scroll-direction-amount\"><code>SCROLL &lt;direction&gt; &lt;amount&gt;</code></h4>\n<p>Scroll the page in a specified direction.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">SCROLL <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">direction</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">amount</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>direction</code> - Direction to scroll: <code>UP</code>, <code>DOWN</code>, <code>LEFT</code>, <code>RIGHT</code>\n- <code>amount</code> - Number of pixels to scroll (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">SCROLL DOWN <span class=\"hljs-number\">500</span>\nSCROLL UP <span class=\"hljs-number\">200</span>\nSCROLL <span class=\"hljs-attribute\">LEFT</span> <span class=\"hljs-number\">100</span>\nSCROLL <span class=\"hljs-attribute\">RIGHT</span> <span class=\"hljs-number\">300</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Smooth scrolling animation\n- Useful for infinite scroll pages\n- Amount can be larger than viewport</p>\n<hr>\n<h4 id=\"move-x-y\"><code>MOVE &lt;x&gt; &lt;y&gt;</code></h4>\n<p>Move mouse cursor to specific coordinates.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">MOVE <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">x</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">y</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>x</code> - X coordinate in pixels (number)\n- <code>y</code> - Y coordinate in pixels (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">MOVE <span class=\"hljs-number\">200</span> <span class=\"hljs-number\">100</span>\nMOVE <span class=\"hljs-number\">500</span> <span class=\"hljs-number\">400</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Triggers hover effects\n- Useful for testing mouseover interactions\n- Does not click, only moves cursor</p>\n<hr>\n<h4 id=\"drag-x1-y1-x2-y2\"><code>DRAG &lt;x1&gt; &lt;y1&gt; &lt;x2&gt; &lt;y2&gt;</code></h4>\n<p>Drag from one point to another.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">DRAG <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">x1</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">y1</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">x2</span>&gt;</span> <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">y2</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>x1</code>, <code>y1</code> - Starting coordinates (numbers)\n- <code>x2</code>, <code>y2</code> - Ending coordinates (numbers)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">DRAG 100 100 500 300\nDRAG 0 200 400 200\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Simulates click, drag, and release\n- Useful for sliders, resizing, reordering\n- Smooth drag animation</p>\n<h3 id=\"keyboard-commands\">⌨️ Keyboard Commands</h3>\n<p>Simulate keyboard input and key presses.</p>\n<h4 id=\"type-text\"><code>TYPE \"&lt;text&gt;\"</code></h4>\n<p>Type text into the currently focused element.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"&lt;text&gt;\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>text</code> - Text to type (string in quotes)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"Hello, World!\"</span>\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"user@example.com\"</span>\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"Password123!\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Requires an input element to be focused\n- Types character by character with realistic timing\n- Supports special characters and Unicode</p>\n<hr>\n<h4 id=\"type-variable\"><code>TYPE $&lt;variable&gt;</code></h4>\n<p>Type the value of a variable.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-variable\">$</span>&lt;variable&gt;\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>variable</code> - Variable name (without quotes)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">SETVAR email = <span class=\"hljs-string\">\"user@example.com\"</span>\nTYPE $email\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Variable must be defined with SETVAR first\n- Variable values are strings\n- Useful for reusable credentials or data</p>\n<hr>\n<h4 id=\"press-key\"><code>PRESS &lt;key&gt;</code></h4>\n<p>Press and release a special key.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">PRESS <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">key</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>key</code> - Key name (see supported keys below)</p>\n<p><strong>Supported Keys:</strong>\n- <code>Tab</code>, <code>Enter</code>, <code>Escape</code>, <code>Space</code>\n- <code>ArrowUp</code>, <code>ArrowDown</code>, <code>ArrowLeft</code>, <code>ArrowRight</code>\n- <code>Delete</code>, <code>Backspace</code>\n- <code>Home</code>, <code>End</code>, <code>PageUp</code>, <code>PageDown</code></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\">PRESS Tab\nPRESS Enter\nPRESS <span class=\"hljs-keyword\">Escape</span>\nPRESS ArrowDown\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Simulates actual key press and release\n- Useful for form navigation and shortcuts\n- Case-sensitive key names</p>\n<hr>\n<h4 id=\"key_down-key\"><code>KEY_DOWN &lt;key&gt;</code></h4>\n<p>Hold down a modifier key.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">KEY_DOWN <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">key</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>key</code> - Modifier key: <code>Shift</code>, <code>Control</code>, <code>Alt</code>, <code>Meta</code></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">KEY_DOWN Shift\nKEY_DOWN Control\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Must be paired with KEY_UP\n- Useful for key combinations\n- Meta key is Cmd on Mac, Windows key on PC</p>\n<hr>\n<h4 id=\"key_up-key\"><code>KEY_UP &lt;key&gt;</code></h4>\n<p>Release a modifier key.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">KEY_UP <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">key</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>key</code> - Modifier key: <code>Shift</code>, <code>Control</code>, <code>Alt</code>, <code>Meta</code></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">KEY_UP Shift\nKEY_UP Control\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Must be paired with KEY_DOWN\n- Releases the specified modifier key\n- Good practice to always release held keys</p>\n<hr>\n<h4 id=\"clear-selector\"><code>CLEAR &lt;selector&gt;</code></h4>\n<p>Clear the content of an input field.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\"><span class=\"hljs-attribute\">CLEAR</span> `&lt;selector&gt;`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for input element (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\"><span class=\"hljs-attribute\">CLEAR</span> `<span class=\"hljs-selector-id\">#search-box</span>`\n<span class=\"hljs-attribute\">CLEAR</span> `<span class=\"hljs-selector-tag\">input</span><span class=\"hljs-selector-attr\">[name=<span class=\"hljs-string\">\"email\"</span>]</span>`\n<span class=\"hljs-attribute\">CLEAR</span> `<span class=\"hljs-selector-class\">.form-input</span><span class=\"hljs-selector-pseudo\">:first</span>-child`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Works with input, textarea elements\n- Faster than selecting all and deleting\n- Triggers appropriate change events</p>\n<hr>\n<h4 id=\"set-selector-value\"><code>SET &lt;selector&gt; \"&lt;value&gt;\"</code></h4>\n<p>Set the value of an input field directly.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">SET <span class=\"hljs-string\">`&lt;selector&gt;`</span> <span class=\"hljs-string\">\"&lt;value&gt;\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector for input element (string in backticks)\n- <code>value</code> - Value to set (string in quotes)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">SET <span class=\"hljs-string\">`#email`</span> <span class=\"hljs-string\">\"user@example.com\"</span>\nSET <span class=\"hljs-string\">`#age`</span> <span class=\"hljs-string\">\"25\"</span>\nSET <span class=\"hljs-string\">`textarea#message`</span> <span class=\"hljs-string\">\"Hello, this is a test message.\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Directly sets value without typing animation\n- Faster than TYPE for long text\n- Triggers change and input events</p>\n<h3 id=\"control-flow-commands\">🔀 Control Flow Commands</h3>\n<p>Add conditional logic and loops to your scripts.</p>\n<h4 id=\"if-exists-selector-then-command\"><code>IF (EXISTS &lt;selector&gt;) THEN &lt;command&gt;</code></h4>\n<p>Execute command if element exists.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">IF (EXISTS `<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">selector</span>&gt;</span>`) THEN <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector to check (string in backticks)\n- <code>command</code> - Command to execute if condition is true</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">IF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept-cookies`</span>\nIF (EXISTS <span class=\"hljs-string\">`#popup-modal`</span>) THEN CLICK <span class=\"hljs-string\">`.close-button`</span>\nIF (EXISTS <span class=\"hljs-string\">`.error-message`</span>) THEN RELOAD\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Checks for element existence at time of execution\n- Does not wait for element to appear\n- Can be combined with ELSE</p>\n<hr>\n<h4 id=\"if-exists-selector-then-command-else-command\"><code>IF (EXISTS &lt;selector&gt;) THEN &lt;command&gt; ELSE &lt;command&gt;</code></h4>\n<p>Execute command based on element existence.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">IF (EXISTS `<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">selector</span>&gt;</span>`) THEN <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span> ELSE <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector to check (string in backticks)\n- First <code>command</code> - Execute if condition is true\n- Second <code>command</code> - Execute if condition is false</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">IF (EXISTS <span class=\"hljs-string\">`.user-menu`</span>) THEN CLICK <span class=\"hljs-string\">`.logout`</span> ELSE CLICK <span class=\"hljs-string\">`.login`</span>\nIF (EXISTS <span class=\"hljs-string\">`.loading`</span>) THEN WAIT <span class=\"hljs-number\">5</span> ELSE CLICK <span class=\"hljs-string\">`#continue`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Exactly one command will be executed\n- Useful for handling different page states\n- Commands must be on same line</p>\n<hr>\n<h4 id=\"if-not-exists-selector-then-command\"><code>IF (NOT EXISTS &lt;selector&gt;) THEN &lt;command&gt;</code></h4>\n<p>Execute command if element does not exist.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">IF (NOT EXISTS `<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">selector</span>&gt;</span>`) THEN <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>selector</code> - CSS selector to check (string in backticks)\n- <code>command</code> - Command to execute if element doesn't exist</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">IF (NOT EXISTS <span class=\"hljs-string\">`.logged-in`</span>) THEN GO /login\nIF (NOT EXISTS <span class=\"hljs-string\">`.results`</span>) THEN CLICK <span class=\"hljs-string\">`#search-button`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Inverse of EXISTS condition\n- Useful for error handling\n- Can check for missing required elements</p>\n<hr>\n<h4 id=\"if-javascript-then-command\"><code>IF (&lt;javascript&gt;) THEN &lt;command&gt;</code></h4>\n<p>Execute command based on JavaScript condition.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">IF (`<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">javascript</span>&gt;</span>`) THEN <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>javascript</code> - JavaScript expression that returns boolean (string in backticks)\n- <code>command</code> - Command to execute if condition is true</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">IF (<span class=\"hljs-string\">`window.innerWidth &lt; 768`</span>) THEN CLICK <span class=\"hljs-string\">`.mobile-menu`</span>\nIF (<span class=\"hljs-string\">`document.readyState === \"complete\"`</span>) THEN CLICK <span class=\"hljs-string\">`#start`</span>\nIF (<span class=\"hljs-string\">`localStorage.getItem(\"user\")`</span>) THEN GO /dashboard\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- JavaScript executes in browser context\n- Must return boolean value\n- Access to all browser APIs and globals</p>\n<hr>\n<h4 id=\"repeat-command-count\"><code>REPEAT (&lt;command&gt;, &lt;count&gt;)</code></h4>\n<p>Repeat a command a specific number of times.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">REPEAT (<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>, <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">count</span>&gt;</span>)\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>command</code> - Command to repeat\n- <code>count</code> - Number of times to repeat (number)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">REPEAT (SCROLL DOWN <span class=\"hljs-number\">300</span>, <span class=\"hljs-number\">5</span>)\nREPEAT (PRESS Tab, <span class=\"hljs-number\">3</span>)\nREPEAT (CLICK `.load-more`, <span class=\"hljs-number\">10</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Executes command exactly count times\n- Useful for pagination, scrolling, navigation\n- No delay between repetitions (add WAIT if needed)</p>\n<hr>\n<h4 id=\"repeat-command-condition\"><code>REPEAT (&lt;command&gt;, &lt;condition&gt;)</code></h4>\n<p>Repeat a command while condition is true.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">REPEAT (<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">command</span>&gt;</span>, `<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">condition</span>&gt;</span>`)\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>command</code> - Command to repeat\n- <code>condition</code> - JavaScript condition to check (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\">REPEAT (SCROLL DOWN <span class=\"hljs-number\">500</span>, `document.querySelector(\".load-more\")`)\nREPEAT (PRESS ArrowDown, `window.scrollY &lt; document.body.scrollHeight`)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Condition checked before each iteration\n- JavaScript condition must return boolean\n- Be careful to avoid infinite loops</p>\n<h3 id=\"variables-and-data\">💾 Variables and Data</h3>\n<p>Store and manipulate data within scripts.</p>\n<h4 id=\"setvar-name-value\"><code>SETVAR &lt;name&gt; = \"&lt;value&gt;\"</code></h4>\n<p>Create or update a variable.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">SETVAR <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">name</span>&gt;</span> = \"<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">value</span>&gt;</span>\"\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>name</code> - Variable name (alphanumeric, underscore)\n- <code>value</code> - Variable value (string in quotes)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-java\"><span class=\"hljs-type\">SETVAR</span> <span class=\"hljs-variable\">username</span> <span class=\"hljs-operator\">=</span> <span class=\"hljs-string\">\"john@example.com\"</span>\n<span class=\"hljs-type\">SETVAR</span> <span class=\"hljs-variable\">password</span> <span class=\"hljs-operator\">=</span> <span class=\"hljs-string\">\"secret123\"</span>\n<span class=\"hljs-type\">SETVAR</span> <span class=\"hljs-variable\">base_url</span> <span class=\"hljs-operator\">=</span> <span class=\"hljs-string\">\"https://api.example.com\"</span>\n<span class=\"hljs-type\">SETVAR</span> <span class=\"hljs-variable\">counter</span> <span class=\"hljs-operator\">=</span> <span class=\"hljs-string\">\"0\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Variables are global within script scope\n- Values are always strings\n- Can be used with TYPE command using $variable syntax</p>\n<hr>\n<h4 id=\"eval-javascript\"><code>EVAL &lt;javascript&gt;</code></h4>\n<p>Execute arbitrary JavaScript code.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-go\">EVAL <span class=\"hljs-string\">`&lt;javascript&gt;`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>javascript</code> - JavaScript code to execute (string in backticks)</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">EVAL `console.<span class=\"hljs-built_in\">log</span>(<span class=\"hljs-string\">\"Script started\"</span>)`\nEVAL `window.<span class=\"hljs-built_in\">scrollTo</span>(<span class=\"hljs-number\">0</span>, <span class=\"hljs-number\">0</span>)`\nEVAL `localStorage.<span class=\"hljs-built_in\">setItem</span>(<span class=\"hljs-string\">\"test\"</span>, <span class=\"hljs-string\">\"value\"</span>)`\nEVAL `document.title = <span class=\"hljs-string\">\"Automated Test\"</span>`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Full access to browser JavaScript APIs\n- Useful for custom logic and debugging\n- Return values are not captured\n- Be careful with security implications</p>\n<h3 id=\"comments-and-documentation\">📝 Comments and Documentation</h3>\n<h4 id=\"comment\"><code># &lt;comment&gt;</code></h4>\n<p>Add comments to scripts for documentation.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># &lt;comment text&gt;</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># This script logs into the application</span>\n<span class=\"hljs-comment\"># Step 1: Navigate to login page</span>\nGO /login\n\n<span class=\"hljs-comment\"># Step 2: Fill credentials</span>\nTYPE <span class=\"hljs-string\">\"user@example.com\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Comments are ignored during execution\n- Useful for documentation and debugging\n- Can appear anywhere in script\n- Supports multi-line documentation blocks</p>\n<h3 id=\"procedures-advanced\">🔧 Procedures (Advanced)</h3>\n<p>Define reusable command sequences.</p>\n<h4 id=\"proc-name-endproc\"><code>PROC &lt;name&gt; ... ENDPROC</code></h4>\n<p>Define a reusable procedure.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">PROC <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">name</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">commands</span>&gt;</span>\nENDPROC\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Parameters:</strong>\n- <code>name</code> - Procedure name (alphanumeric, underscore)\n- <code>commands</code> - Commands to include in procedure</p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">PROC login\n  CLICK <span class=\"hljs-string\">`#email`</span>\n  TYPE $email\n  CLICK <span class=\"hljs-string\">`#password`</span>\n  TYPE $password\n  CLICK <span class=\"hljs-string\">`#submit`</span>\nENDPROC\n\nPROC handle_popups\n  IF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept`</span>\n  IF (EXISTS <span class=\"hljs-string\">`.newsletter-modal`</span>) THEN CLICK <span class=\"hljs-string\">`.close`</span>\nENDPROC\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Procedures must be defined before use\n- Support nested command structures\n- Variables are shared with main script scope</p>\n<hr>\n<h4 id=\"procedure_name\"><code>&lt;procedure_name&gt;</code></h4>\n<p>Call a defined procedure.</p>\n<p><strong>Syntax:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\"><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">procedure_name</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Examples:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"># <span class=\"hljs-keyword\">Define</span> <span class=\"hljs-keyword\">procedure</span> <span class=\"hljs-keyword\">first</span>\nPROC setup\n  GO <span class=\"hljs-operator\">/</span>login\n  WAIT `#form` <span class=\"hljs-number\">5</span>\nENDPROC\n\n# <span class=\"hljs-keyword\">Call</span> <span class=\"hljs-keyword\">procedure</span>\nsetup\nlogin\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Notes:</strong>\n- Procedure must be defined before calling\n- Can be called multiple times\n- No parameters supported (use variables instead)</p>\n<h2 id=\"error-handling-best-practices\">Error Handling Best Practices</h2>\n<h3 id=\"1-always-use-waits\">1. Always Use Waits</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Bad - element might not be ready</span>\nCLICK `<span class=\"hljs-comment\">#button`</span>\n\n<span class=\"hljs-comment\"># Good - wait for element first</span>\nWAIT `<span class=\"hljs-comment\">#button` 5</span>\nCLICK `<span class=\"hljs-comment\">#button`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-handle-optional-elements\">2. Handle Optional Elements</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Check before interacting</span>\nIF (EXISTS <span class=\"hljs-string\">`.popup`</span>) THEN CLICK <span class=\"hljs-string\">`.close`</span>\nIF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept`</span>\n\n<span class=\"hljs-comment\"># Then proceed with main flow</span>\nCLICK <span class=\"hljs-string\">`#main-action`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-use-descriptive-variables\">3. Use Descriptive Variables</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Set up reusable data</span>\nSETVAR admin_email = <span class=\"hljs-string\">\"admin@company.com\"</span>\nSETVAR test_password = <span class=\"hljs-string\">\"TestPass123!\"</span>\nSETVAR staging_url = <span class=\"hljs-string\">\"https://staging.example.com\"</span>\n\n<span class=\"hljs-comment\"># Use throughout script</span>\nGO $staging_url\nTYPE $admin_email\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"4-add-debugging-information\">4. Add Debugging Information</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Log progress</span>\nEVAL `console.log(<span class=\"hljs-string\">\"Starting login process\"</span>)`\nGO /login\n\n<span class=\"hljs-comment\"># Verify page state</span>\nIF (`document.title.includes(<span class=\"hljs-string\">\"Login\"</span>)`) THEN EVAL `console.log(<span class=\"hljs-string\">\"On login page\"</span>)`\n\n<span class=\"hljs-comment\"># Continue with login</span>\nTYPE <span class=\"hljs-variable\">$username</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"common-patterns\">Common Patterns</h2>\n<h3 id=\"login-flow\">Login Flow</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Complete login automation</span>\nSETVAR email = <span class=\"hljs-string\">\"user@example.com\"</span>\nSETVAR password = <span class=\"hljs-string\">\"mypassword\"</span>\n\nGO /login\nWAIT <span class=\"hljs-string\">`#login-form`</span> <span class=\"hljs-number\">5</span>\n\n<span class=\"hljs-comment\"># Handle optional cookie banner</span>\nIF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept-cookies`</span>\n\n<span class=\"hljs-comment\"># Fill and submit form</span>\nCLICK <span class=\"hljs-string\">`#email`</span>\nTYPE $email\nPRESS Tab\nTYPE $password\nCLICK <span class=\"hljs-string\">`button[type=\"submit\"]`</span>\n\n<span class=\"hljs-comment\"># Wait for redirect</span>\nWAIT <span class=\"hljs-string\">`.dashboard`</span> <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"infinite-scroll\">Infinite Scroll</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Load all content with infinite scroll</span>\nGO /products\n\n<span class=\"hljs-comment\"># Scroll and load more content</span>\nREPEAT (SCROLL DOWN 500, `document.querySelector(<span class=\"hljs-string\">\".load-more\"</span>)`)\n\n<span class=\"hljs-comment\"># Alternative: Fixed number of scrolls</span>\nREPEAT (SCROLL DOWN 800, 10)\nWAIT 2\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"form-validation\">Form Validation</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Handle form with validation</span>\nSET <span class=\"hljs-string\">`#email`</span> <span class=\"hljs-string\">\"invalid-email\"</span>\nCLICK <span class=\"hljs-string\">`#submit`</span>\n\n<span class=\"hljs-comment\"># Check for validation error</span>\nIF (EXISTS <span class=\"hljs-string\">`.error-email`</span>) THEN SET <span class=\"hljs-string\">`#email`</span> <span class=\"hljs-string\">\"valid@example.com\"</span>\n\n<span class=\"hljs-comment\"># Retry submission</span>\nCLICK <span class=\"hljs-string\">`#submit`</span>\nWAIT <span class=\"hljs-string\">`.success-message`</span> <span class=\"hljs-number\">5</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"multi-step-process\">Multi-step Process</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Complex multi-step workflow</span>\nPROC navigate_to_step\n  CLICK <span class=\"hljs-string\">`.next-button`</span>\n  WAIT <span class=\"hljs-string\">`.step-content`</span> <span class=\"hljs-number\">5</span>\nENDPROC\n\n<span class=\"hljs-comment\"># Step 1</span>\nWAIT <span class=\"hljs-string\">`.step-1`</span> <span class=\"hljs-number\">5</span>\nSET <span class=\"hljs-string\">`#name`</span> <span class=\"hljs-string\">\"John Doe\"</span>\nnavigate_to_step\n\n<span class=\"hljs-comment\"># Step 2</span>\nSET <span class=\"hljs-string\">`#email`</span> <span class=\"hljs-string\">\"john@example.com\"</span>\nnavigate_to_step\n\n<span class=\"hljs-comment\"># Step 3</span>\nCLICK <span class=\"hljs-string\">`#submit-final`</span>\nWAIT <span class=\"hljs-string\">`.confirmation`</span> <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"integration-with-crawl4ai\">Integration with Crawl4AI</h2>\n<p>Use C4A-Script with Crawl4AI for dynamic content interaction:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Define interaction script</span>\nscript = <span class=\"hljs-string\">\"\"\"\n# Handle dynamic content loading\nWAIT `.content` 5\nIF (EXISTS `.load-more-button`) THEN CLICK `.load-more-button`\nWAIT `.additional-content` 5\n\n# Accept cookies if needed\nIF (EXISTS `.cookie-banner`) THEN CLICK `.accept-all`\n\"\"\"</span>\n\nconfig = CrawlerRunConfig(\n    c4a_script=script,\n    wait_for=<span class=\"hljs-string\">\".content\"</span>,\n    screenshot=<span class=\"hljs-literal\">True</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n    <span class=\"hljs-built_in\">print</span>(result.markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This reference covers all available C4A-Script commands and patterns. For interactive learning, try the <a href=\"../examples/c4a_script/tutorial/\">tutorial</a> or <a href=\"https://docs.crawl4ai.com/c4a-script/demo\">live demo</a>.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/crawl-result/",
    "title": "CrawlResult - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawlresult-reference\"><code>CrawlResult</code> Reference</h1>\n<p>The <strong><code>CrawlResult</code></strong> class encapsulates everything returned after a single crawl operation. It provides the <strong>raw or processed content</strong>, details on links and media, plus optional metadata (like screenshots, PDFs, or extracted JSON).</p>\n<p><strong>Location</strong>: <code>crawl4ai/crawler/models.py</code> (for reference)</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CrawlResult</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    url: <span class=\"hljs-built_in\">str</span>\n    html: <span class=\"hljs-built_in\">str</span>\n    success: <span class=\"hljs-built_in\">bool</span>\n    cleaned_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    fit_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>  <span class=\"hljs-comment\"># Preprocessed HTML optimized for extraction</span>\n    media: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    links: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    downloaded_files: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>]] = <span class=\"hljs-literal\">None</span>\n    screenshot: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    pdf : <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">bytes</span>] = <span class=\"hljs-literal\">None</span>\n    mhtml: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    markdown: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, MarkdownGenerationResult]] = <span class=\"hljs-literal\">None</span>\n    extracted_content: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    metadata: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    error_message: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    session_id: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    response_headers: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    status_code: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">int</span>] = <span class=\"hljs-literal\">None</span>\n    redirected_status_code: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">int</span>] = <span class=\"hljs-literal\">None</span>\n    ssl_certificate: <span class=\"hljs-type\">Optional</span>[SSLCertificate] = <span class=\"hljs-literal\">None</span>\n    dispatch_result: <span class=\"hljs-type\">Optional</span>[DispatchResult] = <span class=\"hljs-literal\">None</span>\n    ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Below is a <strong>field-by-field</strong> explanation and possible usage patterns.</p>\n<hr>\n<h2 id=\"1-basic-crawl-info\">1. Basic Crawl Info</h2>\n<h3 id=\"11-url-str\">1.1 <strong><code>url</code></strong> <em>(str)</em></h3>\n<p><strong>What</strong>: The final crawled URL (after any redirects).<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-built_in\">print</span>(result.url)  <span class=\"hljs-comment\"># e.g., \"https://example.com/\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"12-success-bool\">1.2 <strong><code>success</code></strong> <em>(bool)</em></h3>\n<p><strong>What</strong>: <code>True</code> if the crawl pipeline ended without major errors; <code>False</code> otherwise.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"13-status_code-optionalint\">1.3 <strong><code>status_code</code></strong> <em>(Optional[int])</em></h3>\n<p><strong>What</strong>: The page's HTTP status code (e.g., 200, 404). When the page was reached via redirect, this is the status code of the <strong>first</strong> response in the redirect chain (e.g., 301 or 302).\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.status_code == 404:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Page not found!\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"14-redirected_status_code-optionalint\">1.4 <strong><code>redirected_status_code</code></strong> <em>(Optional[int])</em></h3>\n<p><strong>What</strong>: The HTTP status code of the <strong>final</strong> redirect destination. For a 302→200 redirect, <code>status_code</code> is 302 and <code>redirected_status_code</code> is 200. <code>None</code> for non-HTTP requests (raw HTML, local files).\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.status_code <span class=\"hljs-keyword\">in</span> (<span class=\"hljs-number\">301</span>, <span class=\"hljs-number\">302</span>) <span class=\"hljs-keyword\">and</span> result.redirected_status_code == <span class=\"hljs-number\">200</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Redirected to <span class=\"hljs-subst\">{result.redirected_url}</span> (OK)\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"15-error_message-optionalstr\">1.5 <strong><code>error_message</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: If <code>success=False</code>, a textual description of the failure.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> not result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"15-session_id-optionalstr\">1.5 <strong><code>session_id</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: The ID used for reusing a browser context across multiple calls.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># If you used session_id=\"login_session\" in CrawlerRunConfig, see it here:</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Session:\"</span>, result.session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"16-response_headers-optionaldict\">1.6 <strong><code>response_headers</code></strong> <em>(Optional[dict])</em></h3>\n<p><strong>What</strong>: Final HTTP response headers.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.response_headers</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Server:\"</span>, result.response_headers.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"Server\"</span>, <span class=\"hljs-string\">\"Unknown\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"17-ssl_certificate-optionalsslcertificate\">1.7 <strong><code>ssl_certificate</code></strong> <em>(Optional[SSLCertificate])</em></h3>\n<p><strong>What</strong>: If <code>fetch_ssl_certificate=True</code> in your CrawlerRunConfig, <strong><code>result.ssl_certificate</code></strong> contains a  <a href=\"../../advanced/ssl-certificate/\"><strong><code>SSLCertificate</code></strong></a> object describing the site's certificate. You can export the cert in multiple formats (PEM/DER/JSON) or access its properties like <code>issuer</code>, \n <code>subject</code>, <code>valid_from</code>, <code>valid_until</code>, etc. \n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.ssl_certificate:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Issuer:\"</span>, result.ssl_certificate.issuer)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"2-raw-cleaned-content\">2. Raw / Cleaned Content</h2>\n<h3 id=\"21-html-str\">2.1 <strong><code>html</code></strong> <em>(str)</em></h3>\n<p><strong>What</strong>: The <strong>original</strong> unmodified HTML from the final page load.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Possibly large</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.html))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"22-cleaned_html-optionalstr\">2.2 <strong><code>cleaned_html</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: A sanitized HTML version—scripts, styles, or excluded tags are removed based on your <code>CrawlerRunConfig</code>.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-built_in\">print</span>(result.cleaned_html[:500])  <span class=\"hljs-comment\"># Show a snippet</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"3-markdown-fields\">3. Markdown Fields</h2>\n<h3 id=\"31-the-markdown-generation-approach\">3.1 The Markdown Generation Approach</h3>\n<p>Crawl4AI can convert HTML→Markdown, optionally including:</p>\n<ul>\n<li><strong>Raw</strong> markdown  </li>\n<li><strong>Links as citations</strong> (with a references section)  </li>\n<li><strong>Fit</strong> markdown if a <strong>content filter</strong> is used (like Pruning or BM25)</li>\n</ul>\n<p><strong><code>MarkdownGenerationResult</code></strong> includes:\n- <strong><code>raw_markdown</code></strong> <em>(str)</em>: The full HTML→Markdown conversion.<br>\n- <strong><code>markdown_with_citations</code></strong> <em>(str)</em>: Same markdown, but with link references as academic-style citations.<br>\n- <strong><code>references_markdown</code></strong> <em>(str)</em>: The reference list or footnotes at the end.<br>\n- <strong><code>fit_markdown</code></strong> <em>(Optional[str])</em>: If content filtering (Pruning/BM25) was applied, the filtered \"fit\" text.<br>\n- <strong><code>fit_html</code></strong> <em>(Optional[str])</em>: The HTML that led to <code>fit_markdown</code>.</p>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.markdown:\n    md_res = result.markdown\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw MD:\"</span>, md_res.raw_markdown[:300])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Citations MD:\"</span>, md_res.markdown_with_citations[:300])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"References:\"</span>, md_res.references_markdown)\n    <span class=\"hljs-keyword\">if</span> md_res.fit_markdown:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Pruned text:\"</span>, md_res.fit_markdown[:300])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"32-markdown-optionalunionstr-markdowngenerationresult\">3.2 <strong><code>markdown</code></strong> <em>(Optional[Union[str, MarkdownGenerationResult]])</em></h3>\n<p><strong>What</strong>: Holds the <code>MarkdownGenerationResult</code>.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\"><span class=\"hljs-built_in\">print</span>(result.markdown.raw_markdown[:<span class=\"hljs-number\">200</span>])\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_html)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Important</strong>: \"Fit\" content (in <code>fit_markdown</code>/<code>fit_html</code>) exists in result.markdown, only if you used a <strong>filter</strong> (like <strong>PruningContentFilter</strong> or <strong>BM25ContentFilter</strong>) within a <code>MarkdownGenerationStrategy</code>.<p></p>\n<hr>\n<h2 id=\"4-media-links\">4. Media &amp; Links</h2>\n<h3 id=\"41-media-dictstr-listdict\">4.1 <strong><code>media</code></strong> <em>(Dict[str, List[Dict]])</em></h3>\n<p><strong>What</strong>: Contains info about discovered images, videos, or audio. Typically keys: <code>\"images\"</code>, <code>\"videos\"</code>, <code>\"audios\"</code>.<br>\n<strong>Common Fields</strong> in each item:</p>\n<ul>\n<li><code>src</code> <em>(str)</em>: Media URL  </li>\n<li><code>alt</code> or <code>title</code> <em>(str)</em>: Descriptive text  </li>\n<li><code>score</code> <em>(float)</em>: Relevance score if the crawler's heuristic found it \"important\"  </li>\n<li><code>desc</code> or <code>description</code> <em>(Optional[str])</em>: Additional context extracted from surrounding text  </li>\n</ul>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">images = result.media.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"images\"</span>, [])\n<span class=\"hljs-keyword\">for</span> img <span class=\"hljs-keyword\">in</span> images:\n    <span class=\"hljs-keyword\">if</span> img.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"score\"</span>, <span class=\"hljs-number\">0</span>) &gt; <span class=\"hljs-number\">5</span>:\n        print(<span class=\"hljs-string\">\"High-value image:\"</span>, img[<span class=\"hljs-string\">\"src\"</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"42-links-dictstr-listdict\">4.2 <strong><code>links</code></strong> <em>(Dict[str, List[Dict]])</em></h3>\n<p><strong>What</strong>: Holds internal and external link data. Usually two keys: <code>\"internal\"</code> and <code>\"external\"</code>.<br>\n<strong>Common Fields</strong>:</p>\n<ul>\n<li><code>href</code> <em>(str)</em>: The link target  </li>\n<li><code>text</code> <em>(str)</em>: Link text  </li>\n<li><code>title</code> <em>(str)</em>: Title attribute  </li>\n<li><code>context</code> <em>(str)</em>: Surrounding text snippet  </li>\n<li><code>domain</code> <em>(str)</em>: If external, the domain</li>\n</ul>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links[<span class=\"hljs-string\">\"internal\"</span>]:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal link to <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span> with text <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'text'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"5-additional-fields\">5. Additional Fields</h2>\n<h3 id=\"51-extracted_content-optionalstr\">5.1 <strong><code>extracted_content</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: If you used <strong><code>extraction_strategy</code></strong> (CSS, LLM, etc.), the structured output (JSON).<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.extracted_content</span>:\n    data = json.<span class=\"hljs-built_in\">loads</span>(result.extracted_content)\n    <span class=\"hljs-built_in\">print</span>(data)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"52-downloaded_files-optionalliststr\">5.2 <strong><code>downloaded_files</code></strong> <em>(Optional[List[str]])</em></h3>\n<p><strong>What</strong>: If <code>accept_downloads=True</code> in your <code>BrowserConfig</code> + <code>downloads_path</code>, lists local file paths for downloaded items.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.downloaded_files:\n    <span class=\"hljs-keyword\">for</span> file_path <span class=\"hljs-keyword\">in</span> result.downloaded_files:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Downloaded:\"</span>, file_path)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"53-screenshot-optionalstr\">5.3 <strong><code>screenshot</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: Base64-encoded screenshot if <code>screenshot=True</code> in <code>CrawlerRunConfig</code>.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> base64\n<span class=\"hljs-keyword\">if</span> result.screenshot:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.png\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(base64.b64decode(result.screenshot))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"54-pdf-optionalbytes\">5.4 <strong><code>pdf</code></strong> <em>(Optional[bytes])</em></h3>\n<p><strong>What</strong>: Raw PDF bytes if <code>pdf=True</code> in <code>CrawlerRunConfig</code>.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.pdf:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.pdf\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(result.pdf)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"55-mhtml-optionalstr\">5.5 <strong><code>mhtml</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: MHTML snapshot of the page if <code>capture_mhtml=True</code> in <code>CrawlerRunConfig</code>. MHTML (MIME HTML) format preserves the entire web page with all its resources (CSS, images, scripts, etc.) in a single file.<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.mhtml:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.mhtml\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(result.mhtml)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"56-metadata-optionaldict\">5.6 <strong><code>metadata</code></strong> <em>(Optional[dict])</em></h3>\n<p><strong>What</strong>: Page-level metadata if discovered (title, description, OG data, etc.).<br>\n<strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.metadata</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Title:\"</span>, result.metadata.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"title\"</span>))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Author:\"</span>, result.metadata.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"author\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"6-dispatch_result-optional\">6. <code>dispatch_result</code> (optional)</h2>\n<p>A <code>DispatchResult</code> object providing additional concurrency and resource usage information when crawling URLs in parallel (e.g., via <code>arun_many()</code> with custom dispatchers). It contains:</p>\n<ul>\n<li><strong><code>task_id</code></strong>: A unique identifier for the parallel task.</li>\n<li><strong><code>memory_usage</code></strong> (float): The memory (in MB) used at the time of completion.</li>\n<li><strong><code>peak_memory</code></strong> (float): The peak memory usage (in MB) recorded during the task's execution.</li>\n<li><strong><code>start_time</code></strong> / <strong><code>end_time</code></strong> (datetime): Time range for this crawling task.</li>\n<li><strong><code>error_message</code></strong> (str): Any dispatcher- or concurrency-related error encountered.</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Example usage:</span>\n<span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n    <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.dispatch_result:\n        dr = result.dispatch_result\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"URL: <span class=\"hljs-subst\">{result.url}</span>, Task ID: <span class=\"hljs-subst\">{dr.task_id}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Memory: <span class=\"hljs-subst\">{dr.memory_usage:<span class=\"hljs-number\">.1</span>f}</span> MB (Peak: <span class=\"hljs-subst\">{dr.peak_memory:<span class=\"hljs-number\">.1</span>f}</span> MB)\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Duration: <span class=\"hljs-subst\">{dr.end_time - dr.start_time}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<blockquote>\n<p><strong>Note</strong>: This field is typically populated when using <code>arun_many(...)</code> alongside a <strong>dispatcher</strong> (e.g., <code>MemoryAdaptiveDispatcher</code> or <code>SemaphoreDispatcher</code>). If no concurrency or dispatcher is used, <code>dispatch_result</code> may remain <code>None</code>. </p>\n</blockquote>\n<hr>\n<h2 id=\"7-network-requests-console-messages\">7. Network Requests &amp; Console Messages</h2>\n<p>When you enable network and console message capturing in <code>CrawlerRunConfig</code> using <code>capture_network_requests=True</code> and <code>capture_console_messages=True</code>, the <code>CrawlResult</code> will include these fields:</p>\n<h3 id=\"71-network_requests-optionallistdictstr-any\">7.1 <strong><code>network_requests</code></strong> <em>(Optional[List[Dict[str, Any]]])</em></h3>\n<p><strong>What</strong>: A list of dictionaries containing information about all network requests, responses, and failures captured during the crawl.\n<strong>Structure</strong>:\n- Each item has an <code>event_type</code> field that can be <code>\"request\"</code>, <code>\"response\"</code>, or <code>\"request_failed\"</code>.\n- Request events include <code>url</code>, <code>method</code>, <code>headers</code>, <code>post_data</code>, <code>resource_type</code>, and <code>is_navigation_request</code>.\n- Response events include <code>url</code>, <code>status</code>, <code>status_text</code>, <code>headers</code>, and <code>request_timing</code>.\n- Failed request events include <code>url</code>, <code>method</code>, <code>resource_type</code>, and <code>failure_text</code>.\n- All events include a <code>timestamp</code> field.</p>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.network_requests:\n    <span class=\"hljs-comment\"># Count different types of events</span>\n    requests = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request\"</span>]\n    responses = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"response\"</span>]\n    failures = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request_failed\"</span>]\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(requests)}</span> requests, <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(responses)}</span> responses, and <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(failures)}</span> failures\"</span>)\n\n    <span class=\"hljs-comment\"># Analyze API calls</span>\n    api_calls = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> requests <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"api\"</span> <span class=\"hljs-keyword\">in</span> r.get(<span class=\"hljs-string\">\"url\"</span>, <span class=\"hljs-string\">\"\"</span>)]\n\n    <span class=\"hljs-comment\"># Identify failed resources</span>\n    <span class=\"hljs-keyword\">for</span> failure <span class=\"hljs-keyword\">in</span> failures:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to load: <span class=\"hljs-subst\">{failure.get(<span class=\"hljs-string\">'url'</span>)}</span> - <span class=\"hljs-subst\">{failure.get(<span class=\"hljs-string\">'failure_text'</span>)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"72-console_messages-optionallistdictstr-any\">7.2 <strong><code>console_messages</code></strong> <em>(Optional[List[Dict[str, Any]]])</em></h3>\n<p><strong>What</strong>: A list of dictionaries containing all browser console messages captured during the crawl.\n<strong>Structure</strong>:\n- Each item has a <code>type</code> field indicating the message type (e.g., <code>\"log\"</code>, <code>\"error\"</code>, <code>\"warning\"</code>, etc.).\n- The <code>text</code> field contains the actual message text.\n- Some messages include <code>location</code> information (URL, line, column).\n- All messages include a <code>timestamp</code> field.</p>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.console_messages:\n    <span class=\"hljs-comment\"># Count messages by type</span>\n    message_types = {}\n    <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n        msg_type = msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)\n        message_types[msg_type] = message_types.get(msg_type, <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Message type counts: <span class=\"hljs-subst\">{message_types}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Display errors (which are usually most important)</span>\n    <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n        <span class=\"hljs-keyword\">if</span> msg.get(<span class=\"hljs-string\">\"type\"</span>) == <span class=\"hljs-string\">\"error\"</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{msg.get(<span class=\"hljs-string\">'text'</span>)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>These fields provide deep visibility into the page's network activity and browser console, which is invaluable for debugging, security analysis, and understanding complex web applications.</p>\n<p>For more details on network and console capturing, see the <a href=\"../../advanced/network-console-capture/\">Network &amp; Console Capture documentation</a>.</p>\n<hr>\n<h2 id=\"8-example-accessing-everything\">8. Example: Accessing Everything</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">handle_result</span>(<span class=\"hljs-params\">result: CrawlResult</span>):\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl error:\"</span>, result.error_message)\n        <span class=\"hljs-keyword\">return</span>\n\n    <span class=\"hljs-comment\"># Basic info</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawled URL:\"</span>, result.url)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Status code:\"</span>, result.status_code)\n\n    <span class=\"hljs-comment\"># HTML</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Original HTML size:\"</span>, <span class=\"hljs-built_in\">len</span>(result.html))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML size:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html <span class=\"hljs-keyword\">or</span> <span class=\"hljs-string\">\"\"</span>))\n\n    <span class=\"hljs-comment\"># Markdown output</span>\n    <span class=\"hljs-keyword\">if</span> result.markdown:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown:\"</span>, result.markdown.raw_markdown[:<span class=\"hljs-number\">300</span>])\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Citations Markdown:\"</span>, result.markdown.markdown_with_citations[:<span class=\"hljs-number\">300</span>])\n        <span class=\"hljs-keyword\">if</span> result.markdown.fit_markdown:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown:\"</span>, result.markdown.fit_markdown[:<span class=\"hljs-number\">200</span>])\n\n    <span class=\"hljs-comment\"># Media &amp; Links</span>\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"images\"</span> <span class=\"hljs-keyword\">in</span> result.media:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Image count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.media[<span class=\"hljs-string\">\"images\"</span>]))\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"internal\"</span> <span class=\"hljs-keyword\">in</span> result.links:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal link count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links[<span class=\"hljs-string\">\"internal\"</span>]))\n\n    <span class=\"hljs-comment\"># Extraction strategy result</span>\n    <span class=\"hljs-keyword\">if</span> result.extracted_content:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Structured data:\"</span>, result.extracted_content)\n\n    <span class=\"hljs-comment\"># Screenshot/PDF/MHTML</span>\n    <span class=\"hljs-keyword\">if</span> result.screenshot:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n    <span class=\"hljs-keyword\">if</span> result.pdf:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"PDF bytes length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.pdf))\n    <span class=\"hljs-keyword\">if</span> result.mhtml:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"MHTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.mhtml))\n\n    <span class=\"hljs-comment\"># Network and console capturing</span>\n    <span class=\"hljs-keyword\">if</span> result.network_requests:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Network requests captured: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.network_requests)}</span>\"</span>)\n        <span class=\"hljs-comment\"># Analyze request types</span>\n        req_types = {}\n        <span class=\"hljs-keyword\">for</span> req <span class=\"hljs-keyword\">in</span> result.network_requests:\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"resource_type\"</span> <span class=\"hljs-keyword\">in</span> req:\n                req_types[req[<span class=\"hljs-string\">\"resource_type\"</span>]] = req_types.get(req[<span class=\"hljs-string\">\"resource_type\"</span>], <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Resource types: <span class=\"hljs-subst\">{req_types}</span>\"</span>)\n\n    <span class=\"hljs-keyword\">if</span> result.console_messages:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Console messages captured: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.console_messages)}</span>\"</span>)\n        <span class=\"hljs-comment\"># Count by message type</span>\n        msg_types = {}\n        <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n            msg_types[msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)] = msg_types.get(msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>), <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Message types: <span class=\"hljs-subst\">{msg_types}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"9-key-points-future\">9. Key Points &amp; Future</h2>\n<p>1. <strong>Deprecated legacy properties of CrawlResult</strong><br>\n   - <code>markdown_v2</code> - Removed in v0.5 and now raises <code>AttributeError</code>. Use <code>result.markdown</code> instead.\n   - <code>fit_markdown</code> and <code>fit_html</code> - No longer top-level properties in v0.5. Use <code>result.markdown.fit_markdown</code> and <code>result.markdown.fit_html</code>.</p>\n<p>2. <strong>Fit Content</strong><br>\n   - <strong><code>fit_markdown</code></strong> and <strong><code>fit_html</code></strong> appear in MarkdownGenerationResult, only if you used a content filter (like <strong>PruningContentFilter</strong> or <strong>BM25ContentFilter</strong>) inside your <strong>MarkdownGenerationStrategy</strong> or set them directly.<br>\n   - If no filter is used, they remain <code>None</code>.</p>\n<p>3. <strong>References &amp; Citations</strong><br>\n   - If you enable link citations in your <code>DefaultMarkdownGenerator</code> (<code>options={\"citations\": True}</code>), you’ll see <code>markdown_with_citations</code> plus a <strong><code>references_markdown</code></strong> block. This helps large language models or academic-like referencing.</p>\n<p>4. <strong>Links &amp; Media</strong><br>\n   - <code>links[\"internal\"]</code> and <code>links[\"external\"]</code> group discovered anchors by domain.<br>\n   - <code>media[\"images\"]</code> / <code>[\"videos\"]</code> / <code>[\"audios\"]</code> store extracted media elements with optional scoring or context.</p>\n<p>5. <strong>Error Cases</strong><br>\n   - If <code>success=False</code>, check <code>error_message</code> (e.g., timeouts, invalid URLs).<br>\n   - <code>status_code</code> might be <code>None</code> if we failed before an HTTP response.</p>\n<p>Use <strong><code>CrawlResult</code></strong> to glean all final outputs and feed them into your data pipelines, AI models, or archives. With the synergy of a properly configured <strong>BrowserConfig</strong> and <strong>CrawlerRunConfig</strong>, the crawler can produce robust, structured results here in <strong><code>CrawlResult</code></strong>.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/digest/",
    "title": "digest() - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"digest\">digest()</h1>\n<p>The <code>digest()</code> method is the primary interface for adaptive web crawling. It intelligently crawls websites starting from a given URL, guided by a query, and automatically determines when sufficient information has been gathered.</p>\n<h2 id=\"method-signature\">Method Signature</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">digest</span>(<span class=\"hljs-params\">\n    start_url: <span class=\"hljs-built_in\">str</span>,\n    query: <span class=\"hljs-built_in\">str</span>,\n    resume_from: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, Path]] = <span class=\"hljs-literal\">None</span>\n</span>) -&gt; CrawlState\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"parameters\">Parameters</h2>\n<h3 id=\"start_url\">start_url</h3>\n<ul>\n<li><strong>Type</strong>: <code>str</code></li>\n<li><strong>Required</strong>: Yes</li>\n<li><strong>Description</strong>: The starting URL for the crawl. This should be a valid HTTP/HTTPS URL that serves as the entry point for information gathering.</li>\n</ul>\n<h3 id=\"query\">query</h3>\n<ul>\n<li><strong>Type</strong>: <code>str</code>  </li>\n<li><strong>Required</strong>: Yes</li>\n<li><strong>Description</strong>: The search query that guides the crawling process. This should contain key terms related to the information you're seeking. The crawler uses this to evaluate relevance and determine which links to follow.</li>\n</ul>\n<h3 id=\"resume_from\">resume_from</h3>\n<ul>\n<li><strong>Type</strong>: <code>Optional[Union[str, Path]]</code></li>\n<li><strong>Default</strong>: <code>None</code></li>\n<li><strong>Description</strong>: Path to a previously saved crawl state file. When provided, the crawler resumes from the saved state instead of starting fresh.</li>\n</ul>\n<h2 id=\"return-value\">Return Value</h2>\n<p>Returns a <code>CrawlState</code> object containing:</p>\n<ul>\n<li><strong>crawled_urls</strong> (<code>Set[str]</code>): All URLs that have been crawled</li>\n<li><strong>knowledge_base</strong> (<code>List[CrawlResult]</code>): Collection of crawled pages with content</li>\n<li><strong>pending_links</strong> (<code>List[Link]</code>): Links discovered but not yet crawled</li>\n<li><strong>metrics</strong> (<code>Dict[str, float]</code>): Performance and quality metrics</li>\n<li><strong>query</strong> (<code>str</code>): The original query</li>\n<li>Additional statistical information for scoring</li>\n</ul>\n<h2 id=\"how-it-works\">How It Works</h2>\n<p>The <code>digest()</code> method implements an intelligent crawling algorithm:</p>\n<ol>\n<li><strong>Initial Crawl</strong>: Starts from the provided URL</li>\n<li><strong>Link Analysis</strong>: Evaluates all discovered links for relevance</li>\n<li><strong>Scoring</strong>: Uses three metrics to assess information sufficiency:</li>\n<li><strong>Coverage</strong>: How well the query terms are covered</li>\n<li><strong>Consistency</strong>: Information coherence across pages</li>\n<li><strong>Saturation</strong>: Diminishing returns detection</li>\n<li><strong>Adaptive Selection</strong>: Chooses the most promising links to follow</li>\n<li><strong>Stopping Decision</strong>: Automatically stops when confidence threshold is reached</li>\n</ol>\n<h2 id=\"examples\">Examples</h2>\n<h3 id=\"basic-usage\">Basic Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    adaptive = AdaptiveCrawler(crawler)\n\n    state = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n        start_url=<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>,\n        query=<span class=\"hljs-string\">\"async await context managers\"</span>\n    )\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(state.crawled_urls)}</span> pages\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Confidence: <span class=\"hljs-subst\">{adaptive.confidence:<span class=\"hljs-number\">.0</span>%}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"with-configuration\">With Configuration</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = AdaptiveConfig(\n    confidence_threshold=0.9,  <span class=\"hljs-comment\"># Require high confidence</span>\n    max_pages=30,             <span class=\"hljs-comment\"># Allow more pages</span>\n    top_k_links=3             <span class=\"hljs-comment\"># Follow top 3 links per page</span>\n)\n\nadaptive = AdaptiveCrawler(crawler, config=config)\n\nstate = await adaptive.digest(\n    start_url=<span class=\"hljs-string\">\"https://api.example.com/docs\"</span>,\n    query=<span class=\"hljs-string\">\"authentication endpoints rate limits\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"resuming-a-previous-crawl\">Resuming a Previous Crawl</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># First crawl - may be interrupted</span>\nstate1 <span class=\"hljs-punctuation\">=</span> await adaptive.digest<span class=\"hljs-punctuation\">(</span>\n    start_url<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"machine learning algorithms\"</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Save state (if not auto-saved)</span>\nstate1.save<span class=\"hljs-punctuation\">(</span><span class=\"hljs-string\">\"ml_crawl_state.json\"</span><span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Later, resume from saved state</span>\nstate2 <span class=\"hljs-punctuation\">=</span> await adaptive.digest<span class=\"hljs-punctuation\">(</span>\n    start_url<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"machine learning algorithms\"</span>,\n    resume_from<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"ml_crawl_state.json\"</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"with-progress-monitoring\">With Progress Monitoring</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">state = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n    start_url=<span class=\"hljs-string\">\"https://docs.example.com\"</span>,\n    query=<span class=\"hljs-string\">\"api reference\"</span>\n)\n\n<span class=\"hljs-comment\"># Monitor progress</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Pages crawled: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(state.crawled_urls)}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"New terms discovered: <span class=\"hljs-subst\">{state.new_terms_history}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Final confidence: <span class=\"hljs-subst\">{adaptive.confidence:<span class=\"hljs-number\">.2</span>%}</span>\"</span>)\n\n<span class=\"hljs-comment\"># View detailed statistics</span>\nadaptive.print_stats(detailed=<span class=\"hljs-literal\">True</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"query-best-practices\">Query Best Practices</h2>\n<ol>\n<li>\n<p><strong>Be Specific</strong>: Use descriptive terms that appear in target content\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Good</span>\n<span class=\"hljs-attr\">query</span> = <span class=\"hljs-string\">\"python async context managers implementation\"</span>\n\n<span class=\"hljs-comment\"># Too broad</span>\n<span class=\"hljs-attr\">query</span> = <span class=\"hljs-string\">\"python programming\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Include Key Terms</strong>: Add technical terms you expect to find\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">query</span> = <span class=\"hljs-string\">\"oauth2 jwt refresh tokens authorization\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Multiple Concepts</strong>: Combine related concepts for comprehensive coverage\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">query</span> = <span class=\"hljs-string\">\"rest api pagination sorting filtering\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n<h2 id=\"performance-considerations\">Performance Considerations</h2>\n<ul>\n<li><strong>Initial URL</strong>: Choose a page with good navigation (e.g., documentation index)</li>\n<li><strong>Query Length</strong>: 3-8 terms typically work best</li>\n<li><strong>Link Density</strong>: Sites with clear navigation crawl more efficiently</li>\n<li><strong>Caching</strong>: Enable caching for repeated crawls of the same domain</li>\n</ul>\n<h2 id=\"error-handling\">Error Handling</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">try</span>:\n    state = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n        start_url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        query=<span class=\"hljs-string\">\"search terms\"</span>\n    )\n<span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{e}</span>\"</span>)\n    <span class=\"hljs-comment\"># State is auto-saved if save_state=True in config</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"stopping-conditions\">Stopping Conditions</h2>\n<p>The crawl stops when any of these conditions are met:</p>\n<ol>\n<li><strong>Confidence Threshold</strong>: Reached the configured confidence level</li>\n<li><strong>Page Limit</strong>: Crawled the maximum number of pages</li>\n<li><strong>Diminishing Returns</strong>: Expected information gain below threshold</li>\n<li><strong>No Relevant Links</strong>: No promising links remain to follow</li>\n</ol>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../adaptive-crawler/\">AdaptiveCrawler Class</a></li>\n<li><a href=\"../../core/adaptive-crawling/\">Adaptive Crawling Guide</a></li>\n<li><a href=\"../../core/adaptive-crawling/#configuration-options\">Configuration Options</a></li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/parameters/",
    "title": "Browser, Crawler & LLM Config - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"1-browserconfig-controlling-the-browser\">1. <strong>BrowserConfig</strong> – Controlling the Browser</h1>\n<p><code>BrowserConfig</code> focuses on <strong>how</strong> the browser is launched and behaves. This includes headless mode, proxies, user agents, and other environment tweaks.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nbrowser_cfg = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n    headless=<span class=\"hljs-literal\">True</span>,\n    viewport_width=<span class=\"hljs-number\">1280</span>,\n    viewport_height=<span class=\"hljs-number\">720</span>,\n    proxy_config=<span class=\"hljs-string\">\"http://user:pass@proxy:8080\"</span>,\n    user_agent=<span class=\"hljs-string\">\"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 Chrome/116.0.0.0 Safari/537.36\"</span>,\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"11-parameter-highlights\">1.1 Parameter Highlights</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>browser_type</code></strong></td>\n<td><code>\"chromium\"</code>, <code>\"firefox\"</code>, <code>\"webkit\"</code><br><em>(default: <code>\"chromium\"</code>)</em></td>\n<td>Which browser engine to use. <code>\"chromium\"</code> is typical for many sites, <code>\"firefox\"</code> or <code>\"webkit\"</code> for specialized tests.</td>\n</tr>\n<tr>\n<td><strong><code>headless</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>Headless means no visible UI. <code>False</code> is handy for debugging.</td>\n</tr>\n<tr>\n<td><strong><code>browser_mode</code></strong></td>\n<td><code>str</code> (default: <code>\"dedicated\"</code>)</td>\n<td>How browser is initialized: <code>\"dedicated\"</code> (new instance), <code>\"builtin\"</code> (CDP background), <code>\"custom\"</code> (explicit CDP), <code>\"docker\"</code> (container).</td>\n</tr>\n<tr>\n<td><strong><code>use_managed_browser</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Launch browser via CDP for advanced control. Set automatically based on <code>browser_mode</code>.</td>\n</tr>\n<tr>\n<td><strong><code>cdp_url</code></strong></td>\n<td><code>str</code> (default: <code>None</code>)</td>\n<td>Chrome DevTools Protocol endpoint URL (e.g., <code>\"ws://localhost:9222/devtools/browser/\"</code>). Set automatically based on <code>browser_mode</code>.</td>\n</tr>\n<tr>\n<td><strong><code>debugging_port</code></strong></td>\n<td><code>int</code> (default: <code>9222</code>)</td>\n<td>Port for browser debugging protocol.</td>\n</tr>\n<tr>\n<td><strong><code>host</code></strong></td>\n<td><code>str</code> (default: <code>\"localhost\"</code>)</td>\n<td>Host for browser connection.</td>\n</tr>\n<tr>\n<td><strong><code>viewport_width</code></strong></td>\n<td><code>int</code> (default: <code>1080</code>)</td>\n<td>Initial page width (in px). Useful for testing responsive layouts.</td>\n</tr>\n<tr>\n<td><strong><code>viewport_height</code></strong></td>\n<td><code>int</code> (default: <code>600</code>)</td>\n<td>Initial page height (in px).</td>\n</tr>\n<tr>\n<td><strong><code>viewport</code></strong></td>\n<td><code>dict</code> (default: <code>None</code>)</td>\n<td>Viewport dimensions dict. If set, overrides <code>viewport_width</code> and <code>viewport_height</code>.</td>\n</tr>\n<tr>\n<td><strong><code>device_scale_factor</code></strong></td>\n<td><code>float</code> (default: <code>1.0</code>)</td>\n<td>Device pixel ratio for rendering. Use <code>2.0</code> for Retina-quality screenshots. Higher values produce larger images and use more memory.</td>\n</tr>\n<tr>\n<td><strong><code>proxy</code></strong></td>\n<td><code>str</code> (deprecated)</td>\n<td>Deprecated. Use <code>proxy_config</code> instead. If set, it will be auto-converted internally.</td>\n</tr>\n<tr>\n<td><strong><code>proxy_config</code></strong></td>\n<td><code>ProxyConfig or dict</code> (default: <code>None</code>)</td>\n<td>For advanced or multi-proxy needs, specify <code>ProxyConfig</code> object or dict like <code>{\"server\": \"...\", \"username\": \"...\", \"password\": \"...\"}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>use_persistent_context</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, uses a <strong>persistent</strong> browser context (keep cookies, sessions across runs). Also sets <code>use_managed_browser=True</code>.</td>\n</tr>\n<tr>\n<td><strong><code>user_data_dir</code></strong></td>\n<td><code>str or None</code> (default: <code>None</code>)</td>\n<td>Directory to store user data (profiles, cookies). Must be set if you want permanent sessions.</td>\n</tr>\n<tr>\n<td><strong><code>chrome_channel</code></strong></td>\n<td><code>str</code> (default: <code>\"chromium\"</code>)</td>\n<td>Chrome channel to launch (e.g., \"chrome\", \"msedge\"). Only for <code>browser_type=\"chromium\"</code>. Auto-set to empty for Firefox/WebKit.</td>\n</tr>\n<tr>\n<td><strong><code>channel</code></strong></td>\n<td><code>str</code> (default: <code>\"chromium\"</code>)</td>\n<td>Alias for <code>chrome_channel</code>.</td>\n</tr>\n<tr>\n<td><strong><code>accept_downloads</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Whether to allow file downloads. Requires <code>downloads_path</code> if <code>True</code>.</td>\n</tr>\n<tr>\n<td><strong><code>downloads_path</code></strong></td>\n<td><code>str or None</code> (default: <code>None</code>)</td>\n<td>Directory to store downloaded files.</td>\n</tr>\n<tr>\n<td><strong><code>storage_state</code></strong></td>\n<td><code>str or dict or None</code> (default: <code>None</code>)</td>\n<td>In-memory storage state (cookies, localStorage) to restore browser state.</td>\n</tr>\n<tr>\n<td><strong><code>ignore_https_errors</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>If <code>True</code>, continues despite invalid certificates (common in dev/staging).</td>\n</tr>\n<tr>\n<td><strong><code>java_script_enabled</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>Disable if you want no JS overhead, or if only static content is needed.</td>\n</tr>\n<tr>\n<td><strong><code>sleep_on_close</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Add a small delay when closing browser (can help with cleanup issues).</td>\n</tr>\n<tr>\n<td><strong><code>cookies</code></strong></td>\n<td><code>list</code> (default: <code>[]</code>)</td>\n<td>Pre-set cookies, each a dict like <code>{\"name\": \"session\", \"value\": \"...\", \"url\": \"...\"}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>headers</code></strong></td>\n<td><code>dict</code> (default: <code>{}</code>)</td>\n<td>Extra HTTP headers for every request, e.g. <code>{\"Accept-Language\": \"en-US\"}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent</code></strong></td>\n<td><code>str</code> (default: Chrome-based UA)</td>\n<td>Your custom user agent string.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent_mode</code></strong></td>\n<td><code>str</code> (default: <code>\"\"</code>)</td>\n<td>Set to <code>\"random\"</code> to randomize user agent from a pool (helps with bot detection).</td>\n</tr>\n<tr>\n<td><strong><code>user_agent_generator_config</code></strong></td>\n<td><code>dict</code> (default: <code>{}</code>)</td>\n<td>Configuration dict for user agent generation when <code>user_agent_mode=\"random\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>text_mode</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, tries to disable images/other heavy content for speed.</td>\n</tr>\n<tr>\n<td><strong><code>light_mode</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Disables some background features for performance gains.</td>\n</tr>\n<tr>\n<td><strong><code>avoid_ads</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, blocks requests to common ad/tracker domains (Google Analytics, DoubleClick, Facebook, Hotjar, etc.) at the browser context level.</td>\n</tr>\n<tr>\n<td><strong><code>avoid_css</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, blocks loading of CSS files (<code>.css</code>, <code>.less</code>, <code>.scss</code>, <code>.sass</code>) for faster, leaner crawls when only text content is needed.</td>\n</tr>\n<tr>\n<td><strong><code>extra_args</code></strong></td>\n<td><code>list</code> (default: <code>[]</code>)</td>\n<td>Additional flags for the underlying browser process, e.g. <code>[\"--disable-extensions\"]</code>.</td>\n</tr>\n<tr>\n<td><strong><code>enable_stealth</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Enable playwright-stealth mode to bypass bot detection. Cannot be used with <code>browser_mode=\"builtin\"</code>.</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Tips</strong>:\n- Set <code>headless=False</code> to visually <strong>debug</strong> how pages load or how interactions proceed.<br>\n- If you need <strong>authentication</strong> storage or repeated sessions, consider <code>use_persistent_context=True</code> and specify <code>user_data_dir</code>.<br>\n- For large pages, you might need a bigger <code>viewport_width</code> and <code>viewport_height</code> to handle dynamic content.</p>\n<hr>\n<h1 id=\"2-crawlerrunconfig-controlling-each-crawl\">2. <strong>CrawlerRunConfig</strong> – Controlling Each Crawl</h1>\n<p>While <code>BrowserConfig</code> sets up the <strong>environment</strong>, <code>CrawlerRunConfig</code> details <strong>how</strong> each <strong>crawl operation</strong> should behave: caching, content filtering, link or domain blocking, timeouts, JavaScript code, etc.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\nrun_cfg = CrawlerRunConfig(\n    wait_for=<span class=\"hljs-string\">\"css:.main-content\"</span>,\n    word_count_threshold=<span class=\"hljs-number\">15</span>,\n    excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,\n    stream=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable streaming for arun_many()</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"21-parameter-highlights\">2.1 Parameter Highlights</h2>\n<p>We group them by category. </p>\n<h3 id=\"a-content-processing\">A) <strong>Content Processing</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>word_count_threshold</code></strong></td>\n<td><code>int</code> (default: ~200)</td>\n<td>Skips text blocks below X words. Helps ignore trivial sections.</td>\n</tr>\n<tr>\n<td><strong><code>extraction_strategy</code></strong></td>\n<td><code>ExtractionStrategy</code> (default: None)</td>\n<td>If set, extracts structured data (CSS-based, LLM-based, etc.).</td>\n</tr>\n<tr>\n<td><strong><code>chunking_strategy</code></strong></td>\n<td><code>ChunkingStrategy</code> (default: RegexChunking())</td>\n<td>Strategy to chunk content before extraction. Can be customized for different chunking approaches.</td>\n</tr>\n<tr>\n<td><strong><code>markdown_generator</code></strong></td>\n<td><code>MarkdownGenerationStrategy</code> (None)</td>\n<td>If you want specialized markdown output (citations, filtering, chunking, etc.). Can be customized with options such as <code>content_source</code> parameter to select the HTML input source ('cleaned_html', 'raw_html', or 'fit_html').</td>\n</tr>\n<tr>\n<td><strong><code>css_selector</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>Retains only the part of the page matching this selector. Affects the entire extraction process.</td>\n</tr>\n<tr>\n<td><strong><code>target_elements</code></strong></td>\n<td><code>List[str]</code> (None)</td>\n<td>List of CSS selectors for elements to focus on for markdown generation and data extraction, while still processing the entire page for links, media, etc. Provides more flexibility than <code>css_selector</code>.</td>\n</tr>\n<tr>\n<td><strong><code>excluded_tags</code></strong></td>\n<td><code>list</code> (None)</td>\n<td>Removes entire tags (e.g. <code>[\"script\", \"style\"]</code>).</td>\n</tr>\n<tr>\n<td><strong><code>excluded_selector</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>Like <code>css_selector</code> but to exclude. E.g. <code>\"#ads, .tracker\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>only_text</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, tries to extract text-only content.</td>\n</tr>\n<tr>\n<td><strong><code>prettiify</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, beautifies final HTML (slower, purely cosmetic).</td>\n</tr>\n<tr>\n<td><strong><code>keep_data_attributes</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, preserve <code>data-*</code> attributes in cleaned HTML.</td>\n</tr>\n<tr>\n<td><strong><code>keep_attrs</code></strong></td>\n<td><code>list</code> (default: [])</td>\n<td>List of HTML attributes to keep during processing (e.g., <code>[\"id\", \"class\", \"data-value\"]</code>).</td>\n</tr>\n<tr>\n<td><strong><code>remove_forms</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, remove all <code>&lt;form&gt;</code> elements.</td>\n</tr>\n<tr>\n<td><strong><code>parser_type</code></strong></td>\n<td><code>str</code> (default: \"lxml\")</td>\n<td>HTML parser to use (e.g., \"lxml\", \"html.parser\").</td>\n</tr>\n<tr>\n<td><strong><code>scraping_strategy</code></strong></td>\n<td><code>ContentScrapingStrategy</code> (default: LXMLWebScrapingStrategy())</td>\n<td>Strategy to use for content scraping. Can be customized for different scraping needs (e.g., PDF extraction).</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"b-browser-location-and-identity\">B) <strong>Browser Location and Identity</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>locale</code></strong></td>\n<td><code>str or None</code> (None)</td>\n<td>Browser's locale (e.g., \"en-US\", \"fr-FR\") for language preferences.</td>\n</tr>\n<tr>\n<td><strong><code>timezone_id</code></strong></td>\n<td><code>str or None</code> (None)</td>\n<td>Browser's timezone (e.g., \"America/New_York\", \"Europe/Paris\").</td>\n</tr>\n<tr>\n<td><strong><code>geolocation</code></strong></td>\n<td><code>GeolocationConfig or None</code> (None)</td>\n<td>GPS coordinates configuration. Use <code>GeolocationConfig(latitude=..., longitude=..., accuracy=...)</code>.</td>\n</tr>\n<tr>\n<td><strong><code>fetch_ssl_certificate</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, fetches and includes SSL certificate information in the result.</td>\n</tr>\n<tr>\n<td><strong><code>proxy_config</code></strong></td>\n<td><code>ProxyConfig</code>, <code>list[ProxyConfig]</code>, or <code>None</code> (None)</td>\n<td>Proxy configuration for this specific crawl. Pass a single proxy or an ordered list of proxies to try. See <a href=\"../../advanced/anti-bot-and-fallback/\">Anti-Bot &amp; Fallback</a>.</td>\n</tr>\n<tr>\n<td><strong><code>proxy_rotation_strategy</code></strong></td>\n<td><code>ProxyRotationStrategy</code> (None)</td>\n<td>Strategy for rotating proxies during crawl operations.</td>\n</tr>\n<tr>\n<td><strong><code>max_retries</code></strong></td>\n<td><code>int</code> (0)</td>\n<td>Number of retry rounds when anti-bot blocking is detected. Each round tries all proxies in <code>proxy_config</code>.</td>\n</tr>\n<tr>\n<td><strong><code>fallback_fetch_function</code></strong></td>\n<td><code>async (str) -&gt; str or None</code> (None)</td>\n<td>Async function called as last resort after all retries are exhausted. Takes URL, returns raw HTML. See <a href=\"../../advanced/anti-bot-and-fallback/\">Anti-Bot &amp; Fallback</a>.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"c-caching-session\">C) <strong>Caching &amp; Session</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>cache_mode</code></strong></td>\n<td><code>CacheMode or None</code></td>\n<td>Controls how caching is handled (<code>ENABLED</code>, <code>BYPASS</code>, <code>DISABLED</code>, etc.). If <code>None</code>, typically defaults to <code>ENABLED</code>.</td>\n</tr>\n<tr>\n<td><strong><code>session_id</code></strong></td>\n<td><code>str or None</code></td>\n<td>Assign a unique ID to reuse a single browser session across multiple <code>arun()</code> calls.</td>\n</tr>\n<tr>\n<td><strong><code>bypass_cache</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td><strong>Deprecated.</strong> If <code>True</code>, acts like <code>CacheMode.BYPASS</code>. Use <code>cache_mode</code> instead.</td>\n</tr>\n<tr>\n<td><strong><code>disable_cache</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td><strong>Deprecated.</strong> If <code>True</code>, acts like <code>CacheMode.DISABLED</code>. Use <code>cache_mode</code> instead.</td>\n</tr>\n<tr>\n<td><strong><code>no_cache_read</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td><strong>Deprecated.</strong> If <code>True</code>, acts like <code>CacheMode.WRITE_ONLY</code> (writes cache but never reads). Use <code>cache_mode</code> instead.</td>\n</tr>\n<tr>\n<td><strong><code>no_cache_write</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td><strong>Deprecated.</strong> If <code>True</code>, acts like <code>CacheMode.READ_ONLY</code> (reads cache but never writes). Use <code>cache_mode</code> instead.</td>\n</tr>\n<tr>\n<td><strong><code>shared_data</code></strong></td>\n<td><code>dict or None</code> (None)</td>\n<td>Shared data to be passed between hooks and accessible across crawl operations.</td>\n</tr>\n</tbody>\n</table>\n<p>Use these for controlling whether you read or write from a local content cache. Handy for large batch crawls or repeated site visits.</p>\n<hr>\n<h3 id=\"d-page-navigation-timing\">D) <strong>Page Navigation &amp; Timing</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>wait_until</code></strong></td>\n<td><code>str</code> (domcontentloaded)</td>\n<td>Condition for navigation to \"complete\". Often <code>\"networkidle\"</code> or <code>\"domcontentloaded\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>page_timeout</code></strong></td>\n<td><code>int</code> (60000 ms)</td>\n<td>Timeout for page navigation or JS steps. Increase for slow sites.</td>\n</tr>\n<tr>\n<td><strong><code>wait_for</code></strong></td>\n<td><code>str or None</code></td>\n<td>Wait for a CSS (<code>\"css:selector\"</code>) or JS (<code>\"js:() =&gt; bool\"</code>) condition before content extraction.</td>\n</tr>\n<tr>\n<td><strong><code>wait_for_timeout</code></strong></td>\n<td><code>int or None</code> (None)</td>\n<td>Specific timeout in ms for the <code>wait_for</code> condition. If None, uses <code>page_timeout</code>.</td>\n</tr>\n<tr>\n<td><strong><code>wait_for_images</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Wait for images to load before finishing. Slows down if you only want text.</td>\n</tr>\n<tr>\n<td><strong><code>delay_before_return_html</code></strong></td>\n<td><code>float</code> (0.1)</td>\n<td>Additional pause (seconds) before final HTML is captured. Good for last-second updates.</td>\n</tr>\n<tr>\n<td><strong><code>check_robots_txt</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Whether to check and respect robots.txt rules before crawling. If True, caches robots.txt for efficiency.</td>\n</tr>\n<tr>\n<td><strong><code>mean_delay</code></strong> and <strong><code>max_range</code></strong></td>\n<td><code>float</code> (0.1, 0.3)</td>\n<td>If you call <code>arun_many()</code>, these define random delay intervals between crawls, helping avoid detection or rate limits.</td>\n</tr>\n<tr>\n<td><strong><code>semaphore_count</code></strong></td>\n<td><code>int</code> (5)</td>\n<td>Max concurrency for <code>arun_many()</code>. Increase if you have resources for parallel crawls.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"e-page-interaction\">E) <strong>Page Interaction</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>js_code</code></strong></td>\n<td><code>str or list[str]</code> (None)</td>\n<td>JavaScript to run <strong>after</strong> <code>wait_for</code> and <code>delay_before_return_html</code>, on the fully-loaded page. E.g. <code>\"document.querySelector('button')?.click();\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>js_code_before_wait</code></strong></td>\n<td><code>str or list[str]</code> (None)</td>\n<td>JavaScript to run <strong>before</strong> <code>wait_for</code>. Use for triggering loading that <code>wait_for</code> then checks (e.g. clicking a tab, then waiting for its content).</td>\n</tr>\n<tr>\n<td><strong><code>c4a_script</code></strong></td>\n<td><code>str or list[str]</code> (None)</td>\n<td>C4A script that compiles to JavaScript. Alternative to writing raw JS.</td>\n</tr>\n<tr>\n<td><strong><code>js_only</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, indicates we're reusing an existing session and only applying JS. No full reload.</td>\n</tr>\n<tr>\n<td><strong><code>ignore_body_visibility</code></strong></td>\n<td><code>bool</code> (True)</td>\n<td>Skip checking if <code>&lt;body&gt;</code> is visible. Usually best to keep <code>True</code>.</td>\n</tr>\n<tr>\n<td><strong><code>scan_full_page</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, auto-scroll the page to load dynamic content (infinite scroll).</td>\n</tr>\n<tr>\n<td><strong><code>scroll_delay</code></strong></td>\n<td><code>float</code> (0.2)</td>\n<td>Delay between scroll steps when scanning the full page (<code>scan_full_page=True</code>) or capturing full-page screenshots.</td>\n</tr>\n<tr>\n<td><strong><code>max_scroll_steps</code></strong></td>\n<td><code>int or None</code> (None)</td>\n<td>Maximum number of scroll steps during full page scan. If None, scrolls until entire page is loaded.</td>\n</tr>\n<tr>\n<td><strong><code>process_iframes</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Inlines iframe content for single-page extraction.</td>\n</tr>\n<tr>\n<td><strong><code>flatten_shadow_dom</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Flattens Shadow DOM content into the light DOM before HTML capture. Resolves slots, strips shadow-scoped styles, and force-opens closed shadow roots. Essential for sites built with Web Components (Stencil, Lit, Shoelace, etc.).</td>\n</tr>\n<tr>\n<td><strong><code>remove_overlay_elements</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes potential modals/popups blocking the main content.</td>\n</tr>\n<tr>\n<td><strong><code>remove_consent_popups</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes GDPR/cookie consent popups from known CMP providers (OneTrust, Cookiebot, TrustArc, Quantcast, Didomi, Sourcepoint, FundingChoices, etc.). Tries clicking \"Accept All\" first, then falls back to DOM removal.</td>\n</tr>\n<tr>\n<td><strong><code>simulate_user</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Simulate user interactions (mouse movements) to avoid bot detection.</td>\n</tr>\n<tr>\n<td><strong><code>override_navigator</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Override <code>navigator</code> properties in JS for stealth.</td>\n</tr>\n<tr>\n<td><strong><code>magic</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Automatic handling of popups/consent banners. Experimental.</td>\n</tr>\n<tr>\n<td><strong><code>adjust_viewport_to_content</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Resizes viewport to match page content height.</td>\n</tr>\n</tbody>\n</table>\n<p>If your page is a single-page app with repeated JS updates, set <code>js_only=True</code> in subsequent calls, plus a <code>session_id</code> for reusing the same tab.</p>\n<hr>\n<h3 id=\"f-media-handling\">F) <strong>Media Handling</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>screenshot</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Capture a screenshot (base64) in <code>result.screenshot</code>.</td>\n</tr>\n<tr>\n<td><strong><code>screenshot_wait_for</code></strong></td>\n<td><code>float or None</code></td>\n<td>Extra wait time before the screenshot.</td>\n</tr>\n<tr>\n<td><strong><code>screenshot_height_threshold</code></strong></td>\n<td><code>int</code> (~20000)</td>\n<td>If the page is taller than this, alternate screenshot strategies are used.</td>\n</tr>\n<tr>\n<td><strong><code>force_viewport_screenshot</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, always captures a viewport-only screenshot regardless of page height. Faster and smaller than full-page screenshots.</td>\n</tr>\n<tr>\n<td><strong><code>pdf</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, returns a PDF in <code>result.pdf</code>.</td>\n</tr>\n<tr>\n<td><strong><code>capture_mhtml</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, captures an MHTML snapshot of the page in <code>result.mhtml</code>. MHTML includes all page resources (CSS, images, etc.) in a single file.</td>\n</tr>\n<tr>\n<td><strong><code>image_description_min_word_threshold</code></strong></td>\n<td><code>int</code> (~50)</td>\n<td>Minimum words for an image's alt text or description to be considered valid.</td>\n</tr>\n<tr>\n<td><strong><code>image_score_threshold</code></strong></td>\n<td><code>int</code> (~3)</td>\n<td>Filter out low-scoring images. The crawler scores images by relevance (size, context, etc.).</td>\n</tr>\n<tr>\n<td><strong><code>exclude_external_images</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Exclude images from other domains.</td>\n</tr>\n<tr>\n<td><strong><code>exclude_all_images</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, excludes all images from processing (both internal and external).</td>\n</tr>\n<tr>\n<td><strong><code>table_score_threshold</code></strong></td>\n<td><code>int</code> (7)</td>\n<td>Minimum score threshold for processing a table. Lower values include more tables.</td>\n</tr>\n<tr>\n<td><strong><code>table_extraction</code></strong></td>\n<td><code>TableExtractionStrategy</code> (DefaultTableExtraction)</td>\n<td>Strategy for table extraction. Defaults to DefaultTableExtraction with configured threshold.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"g-linkdomain-handling\">G) <strong>Link/Domain Handling</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>exclude_social_media_domains</code></strong></td>\n<td><code>list</code> (e.g. Facebook/Twitter)</td>\n<td>A default list can be extended. Any link to these domains is removed from final output.</td>\n</tr>\n<tr>\n<td><strong><code>exclude_external_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes all links pointing outside the current domain.</td>\n</tr>\n<tr>\n<td><strong><code>exclude_social_media_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Strips links specifically to social sites (like Facebook or Twitter).</td>\n</tr>\n<tr>\n<td><strong><code>exclude_domains</code></strong></td>\n<td><code>list</code> ([])</td>\n<td>Provide a custom list of domains to exclude (like <code>[\"ads.com\", \"trackers.io\"]</code>).</td>\n</tr>\n<tr>\n<td><strong><code>exclude_internal_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, excludes internal links from the results.</td>\n</tr>\n<tr>\n<td><strong><code>score_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, calculates intrinsic quality scores for all links using URL structure, text quality, and contextual metrics.</td>\n</tr>\n<tr>\n<td><strong><code>preserve_https_for_internal_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, preserves HTTPS scheme for internal links even when the server redirects to HTTP. Useful for security-conscious crawling.</td>\n</tr>\n</tbody>\n</table>\n<p>Use these for link-level content filtering (often to keep crawls “internal” or to remove spammy domains).</p>\n<hr>\n<h3 id=\"h-debug-logging-network-monitoring\">H) <strong>Debug, Logging &amp; Network Monitoring</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>verbose</code></strong></td>\n<td><code>bool</code> (True)</td>\n<td>Prints logs detailing each step of crawling, interactions, or errors.</td>\n</tr>\n<tr>\n<td><strong><code>log_console</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Logs the page's JavaScript console output if you want deeper JS debugging.</td>\n</tr>\n<tr>\n<td><strong><code>capture_network_requests</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, captures network requests made by the page in <code>result.captured_requests</code>.</td>\n</tr>\n<tr>\n<td><strong><code>capture_console_messages</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, captures console messages from the page in <code>result.console_messages</code>.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"i-connection-http-parameters\">I) <strong>Connection &amp; HTTP Parameters</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>method</code></strong></td>\n<td><code>str</code> (\"GET\")</td>\n<td>HTTP method to use when using AsyncHTTPCrawlerStrategy (e.g., \"GET\", \"POST\").</td>\n</tr>\n<tr>\n<td><strong><code>stream</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, enables streaming mode for <code>arun_many()</code> to process URLs as they complete rather than waiting for all.</td>\n</tr>\n<tr>\n<td><strong><code>url</code></strong></td>\n<td><code>str or None</code> (None)</td>\n<td>URL for this specific config. Not typically set directly but used internally for URL-specific configurations.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent</code></strong></td>\n<td><code>str or None</code> (None)</td>\n<td>Custom User-Agent string for this crawl. Can override browser-level user agent.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent_mode</code></strong></td>\n<td><code>str or None</code> (None)</td>\n<td>Set to <code>\"random\"</code> to randomize user agent. Can override browser-level setting.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent_generator_config</code></strong></td>\n<td><code>dict</code> ({})</td>\n<td>Configuration for user agent generation when <code>user_agent_mode=\"random\"</code>.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3 id=\"j-virtual-scroll-configuration\">J) <strong>Virtual Scroll Configuration</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>virtual_scroll_config</code></strong></td>\n<td><code>VirtualScrollConfig or dict</code> (None)</td>\n<td>Configuration for handling virtualized scrolling on sites like Twitter/Instagram where content is replaced rather than appended.</td>\n</tr>\n</tbody>\n</table>\n<p>When sites use virtual scrolling (content replaced as you scroll), use <code>VirtualScrollConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai import VirtualScrollConfig\n\nvirtual_config = VirtualScrollConfig(\n    container_selector=<span class=\"hljs-string\">\"#timeline\"</span>,    <span class=\"hljs-comment\"># CSS selector for scrollable container</span>\n    scroll_count=30,                   <span class=\"hljs-comment\"># Number of times to scroll</span>\n    scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,      <span class=\"hljs-comment\"># How much to scroll: \"container_height\", \"page_height\", or pixels (e.g. 500)</span>\n    wait_after_scroll=0.5             <span class=\"hljs-comment\"># Seconds to wait after each scroll for content to load</span>\n)\n\nconfig = CrawlerRunConfig(\n    virtual_scroll_config=virtual_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>VirtualScrollConfig Parameters:</strong></p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>container_selector</code></strong></td>\n<td><code>str</code> (required)</td>\n<td>CSS selector for the scrollable container (e.g., <code>\"#feed\"</code>, <code>\".timeline\"</code>)</td>\n</tr>\n<tr>\n<td><strong><code>scroll_count</code></strong></td>\n<td><code>int</code> (10)</td>\n<td>Maximum number of scrolls to perform</td>\n</tr>\n<tr>\n<td><strong><code>scroll_by</code></strong></td>\n<td><code>str or int</code> (\"container_height\")</td>\n<td>Scroll amount: <code>\"container_height\"</code>, <code>\"page_height\"</code>, or pixels (e.g., <code>500</code>)</td>\n</tr>\n<tr>\n<td><strong><code>wait_after_scroll</code></strong></td>\n<td><code>float</code> (0.5)</td>\n<td>Time in seconds to wait after each scroll for new content to load</td>\n</tr>\n</tbody>\n</table>\n<p><strong>When to use Virtual Scroll vs scan_full_page:</strong>\n- Use <code>virtual_scroll_config</code> when content is <strong>replaced</strong> during scroll (Twitter, Instagram)\n- Use <code>scan_full_page</code> when content is <strong>appended</strong> during scroll (traditional infinite scroll)</p>\n<p>See <a href=\"../../advanced/virtual-scroll.md\">Virtual Scroll documentation</a> for detailed examples.</p>\n<hr>\n<h3 id=\"k-url-matching-configuration\">K) <strong>URL Matching Configuration</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>url_matcher</code></strong></td>\n<td><code>UrlMatcher</code> (None)</td>\n<td>Pattern(s) to match URLs against. Can be: string (glob), function, or list of mixed types. <strong>None means match ALL URLs</strong></td>\n</tr>\n<tr>\n<td><strong><code>match_mode</code></strong></td>\n<td><code>MatchMode</code> (MatchMode.OR)</td>\n<td>How to combine multiple matchers in a list: <code>MatchMode.OR</code> (any match) or <code>MatchMode.AND</code> (all must match)</td>\n</tr>\n</tbody>\n</table>\n<p>The <code>url_matcher</code> parameter enables URL-specific configurations when used with <code>arun_many()</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, MatchMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.extraction_strategy <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-comment\"># Simple string pattern (glob-style)</span>\npdf_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>,\n    scraping_strategy=PDFContentScrapingStrategy()\n)\n\n<span class=\"hljs-comment\"># Multiple patterns with OR logic (default)</span>\nblog_config = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"*/blog/*\"</span>, <span class=\"hljs-string\">\"*/article/*\"</span>, <span class=\"hljs-string\">\"*/news/*\"</span>],\n    match_mode=MatchMode.OR  <span class=\"hljs-comment\"># Any pattern matches</span>\n)\n\n<span class=\"hljs-comment\"># Function matcher</span>\napi_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'api'</span> <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">or</span> url.endswith(<span class=\"hljs-string\">'.json'</span>),\n    <span class=\"hljs-comment\"># Other settings like extraction_strategy</span>\n)\n\n<span class=\"hljs-comment\"># Mixed: String + Function with AND logic</span>\ncomplex_config = CrawlerRunConfig(\n    url_matcher=[\n        <span class=\"hljs-keyword\">lambda</span> url: url.startswith(<span class=\"hljs-string\">'https://'</span>),  <span class=\"hljs-comment\"># Must be HTTPS</span>\n        <span class=\"hljs-string\">\"*.org/*\"</span>,                               <span class=\"hljs-comment\"># Must be .org domain</span>\n        <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'docs'</span> <span class=\"hljs-keyword\">in</span> url                <span class=\"hljs-comment\"># Must contain 'docs'</span>\n    ],\n    match_mode=MatchMode.AND  <span class=\"hljs-comment\"># ALL conditions must match</span>\n)\n\n<span class=\"hljs-comment\"># Combined patterns and functions with AND logic</span>\nsecure_docs = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"https://*\"</span>, <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'.doc'</span> <span class=\"hljs-keyword\">in</span> url],\n    match_mode=MatchMode.AND  <span class=\"hljs-comment\"># Must be HTTPS AND contain .doc</span>\n)\n\n<span class=\"hljs-comment\"># Default config - matches ALL URLs</span>\ndefault_config = CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher = matches everything</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>UrlMatcher Types:</strong>\n- <strong>None (default)</strong>: When <code>url_matcher</code> is None or not set, the config matches ALL URLs\n- <strong>String patterns</strong>: Glob-style patterns like <code>\"*.pdf\"</code>, <code>\"*/api/*\"</code>, <code>\"https://*.example.com/*\"</code>\n- <strong>Functions</strong>: <code>lambda url: bool</code> - Custom logic for complex matching\n- <strong>Lists</strong>: Mix strings and functions, combined with <code>MatchMode.OR</code> or <code>MatchMode.AND</code></p>\n<p><strong>Important Behavior:</strong>\n- When passing a list of configs to <code>arun_many()</code>, URLs are matched against each config's <code>url_matcher</code> in order. First match wins!\n- If no config matches a URL and there's no default config (one without <code>url_matcher</code>), the URL will fail with \"No matching configuration found\"\n- Always include a default config as the last item if you want to handle all URLs</p>\n<hr>\n<h3 id=\"l-advanced-crawling-features\">L) <strong>Advanced Crawling Features</strong></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>deep_crawl_strategy</code></strong></td>\n<td><code>DeepCrawlStrategy or None</code> (None)</td>\n<td>Strategy for deep/recursive crawling. Enables automatic link following and multi-level site crawling.</td>\n</tr>\n<tr>\n<td><strong><code>link_preview_config</code></strong></td>\n<td><code>LinkPreviewConfig or dict or None</code> (None)</td>\n<td>Configuration for link head extraction and scoring. Fetches and scores link metadata without full page loads.</td>\n</tr>\n<tr>\n<td><strong><code>experimental</code></strong></td>\n<td><code>dict or None</code> (None)</td>\n<td>Dictionary for experimental/beta features not yet integrated into main parameters. Use with caution.</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Deep Crawl Strategy</strong> enables automatic site exploration by following links according to defined rules. Useful for sitemap generation or comprehensive site archiving.</p>\n<p><strong>Link Preview Config</strong> allows efficient link discovery and scoring by fetching only the <code>&lt;head&gt;</code> section of linked pages, enabling smart crawl prioritization without the overhead of full page loads.</p>\n<p><strong>Experimental</strong> parameters are features in beta testing. They may change or be removed in future versions. Check documentation for currently available experimental features.</p>\n<hr>\n<h2 id=\"22-helper-methods\">2.2 Helper Methods</h2>\n<p>Both <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> provide a <code>clone()</code> method to create modified copies:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Create a base configuration</span>\nbase_config = CrawlerRunConfig(\n    cache_mode=CacheMode.ENABLED,\n    word_count_threshold=200\n)\n\n<span class=\"hljs-comment\"># Create variations using clone()</span>\nstream_config = base_config.clone(stream=True)\nno_cache_config = base_config.clone(\n    cache_mode=CacheMode.BYPASS,\n    stream=True\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The <code>clone()</code> method is particularly useful when you need slightly different configurations for different use cases, without modifying the original config.</p>\n<h3 id=\"class-level-defaults-set_defaults-get_defaults-reset_defaults\">Class-Level Defaults (<code>set_defaults</code> / <code>get_defaults</code> / <code>reset_defaults</code>)</h3>\n<p>Both config classes support class-level default overrides. When deploying in a server or cloud context, this eliminates the need to pass the same parameters at every call site.</p>\n<p><strong>Resolution order:</strong> explicit arg &gt; class-level default &gt; hardcoded default</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Set once at application startup</span>\nBrowserConfig.set_defaults(\n    cache_cdp_connection=<span class=\"hljs-literal\">True</span>,\n    cdp_close_delay=<span class=\"hljs-number\">0</span>,\n    create_isolated_context=<span class=\"hljs-literal\">True</span>,\n)\nCrawlerRunConfig.set_defaults(verbose=<span class=\"hljs-literal\">False</span>)\n\n<span class=\"hljs-comment\"># All new instances inherit the class defaults</span>\ncfg1 = BrowserConfig(cdp_url=<span class=\"hljs-string\">\"ws://localhost:9222\"</span>)\n<span class=\"hljs-comment\"># → cache_cdp_connection=True, cdp_close_delay=0</span>\n\ncfg2 = BrowserConfig(cdp_url=<span class=\"hljs-string\">\"ws://localhost:9222\"</span>, cache_cdp_connection=<span class=\"hljs-literal\">False</span>)\n<span class=\"hljs-comment\"># → cache_cdp_connection=False (explicit value wins)</span>\n\n<span class=\"hljs-comment\"># Inspect current defaults</span>\nBrowserConfig.get_defaults()\n<span class=\"hljs-comment\"># → {\"cache_cdp_connection\": True, \"cdp_close_delay\": 0, \"create_isolated_context\": True}</span>\n\n<span class=\"hljs-comment\"># Remove a single default</span>\nBrowserConfig.reset_defaults(<span class=\"hljs-string\">\"cdp_close_delay\"</span>)\n\n<span class=\"hljs-comment\"># Remove all defaults</span>\nBrowserConfig.reset_defaults()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>API Reference:</strong></p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Method</th>\n<th>Signature</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>set_defaults</code></td>\n<td><code>set_defaults(**kwargs)</code></td>\n<td>Set class-level defaults for new instances. Raises <code>ValueError</code> if any key is not a valid <code>__init__</code> parameter.</td>\n</tr>\n<tr>\n<td><code>get_defaults</code></td>\n<td><code>get_defaults() → dict</code></td>\n<td>Return a deep copy of the current class-level defaults.</td>\n</tr>\n<tr>\n<td><code>reset_defaults</code></td>\n<td><code>reset_defaults(*names)</code></td>\n<td>With no args, clears all defaults. With args, removes only the named defaults.</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Notes:</strong>\n- Defaults are independent per class — <code>BrowserConfig.set_defaults()</code> has no effect on <code>CrawlerRunConfig</code>.\n- Mutable values (lists, dicts) are deep-copied on storage and on each instance creation, so instances do not share objects.\n- <code>clone()</code>, <code>dump()</code>/<code>load()</code>, and <code>from_kwargs()</code> all work correctly with class defaults — serialized data is self-contained and independent of the current class defaults.\n- Defaults are stored in memory for the lifetime of the process. They are not persisted to disk.</p>\n<h2 id=\"23-example-usage\">2.3 Example Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Configure the browser</span>\n    browser_cfg = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,\n        viewport_width=<span class=\"hljs-number\">1280</span>,\n        viewport_height=<span class=\"hljs-number\">720</span>,\n        proxy_config=<span class=\"hljs-string\">\"http://user:pass@myproxy:8080\"</span>,\n        text_mode=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># Configure the run</span>\n    run_cfg = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        session_id=<span class=\"hljs-string\">\"my_session\"</span>,\n        css_selector=<span class=\"hljs-string\">\"main.article\"</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"script\"</span>, <span class=\"hljs-string\">\"style\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.article-loaded\"</span>,\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        stream=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/news\"</span>,\n            config=run_cfg\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Final cleaned_html length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot captured (base64, length):\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"24-compliance-ethics\">2.4 Compliance &amp; Ethics</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>check_robots_txt</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>When True, checks and respects robots.txt rules before crawling. Uses efficient caching with SQLite backend.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>User agent string to identify your crawler. Used for robots.txt checking when enabled.</td>\n</tr>\n</tbody>\n</table>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    check_robots_txt<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable robots.txt compliance</span>\n    user_agent<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"MyBot/1.0\"</span>  <span class=\"hljs-comment\"># Identify your crawler</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h1 id=\"3-llmconfig-setting-up-llm-providers\">3. <strong>LLMConfig</strong> - Setting up LLM providers</h1>\n<p>LLMConfig is useful to pass LLM provider config to strategies and functions that rely on LLMs to do extraction, filtering, schema generation etc. Currently it can be used in the following -</p>\n<ol>\n<li>LLMExtractionStrategy</li>\n<li>LLMContentFilter</li>\n<li>JsonCssExtractionStrategy.generate_schema</li>\n<li>JsonXPathExtractionStrategy.generate_schema</li>\n<li>AdaptiveConfig.embedding_llm_config (embedding model for adaptive crawling)</li>\n<li>AdaptiveConfig.query_llm_config (chat completion model for query expansion in adaptive crawling)</li>\n</ol>\n<h2 id=\"31-parameters\">3.1 Parameters</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>provider</code></strong></td>\n<td><code>\"ollama/llama3\",\"groq/llama3-70b-8192\",\"groq/llama3-8b-8192\", \"openai/gpt-4o-mini\" ,\"openai/gpt-4o\",\"openai/o1-mini\",\"openai/o1-preview\",\"openai/o3-mini\",\"openai/o3-mini-high\",\"anthropic/claude-3-haiku-20240307\",\"anthropic/claude-3-opus-20240229\",\"anthropic/claude-3-sonnet-20240229\",\"anthropic/claude-3-5-sonnet-20240620\",\"gemini/gemini-pro\",\"gemini/gemini-1.5-pro\",\"gemini/gemini-2.0-flash\",\"gemini/gemini-2.0-flash-exp\",\"gemini/gemini-2.0-flash-lite-preview-02-05\",\"deepseek/deepseek-chat\"</code><br><em>(default: <code>\"openai/gpt-4o-mini\"</code>)</em></td>\n<td>Which LLM provider to use.</td>\n</tr>\n<tr>\n<td><strong><code>api_token</code></strong></td>\n<td>1.Optional. When not provided explicitly, api_token will be read from environment variables based on provider. For example: If a gemini model is passed as provider then,<code>\"GEMINI_API_KEY\"</code> will be read from environment variables  <br> 2. API token of LLM provider <br> eg: <code>api_token = \"gsk_1ClHGGJ7Lpn4WGybR7vNWGdyb3FY7zXEw3SCiy0BAVM9lL8CQv\"</code> <br> 3. Environment variable - use with prefix \"env:\" <br> eg:<code>api_token = \"env: GROQ_API_KEY\"</code></td>\n<td>API token to use for the given provider</td>\n</tr>\n<tr>\n<td><strong><code>base_url</code></strong></td>\n<td>Optional. Custom API endpoint</td>\n<td>If your provider has a custom endpoint</td>\n</tr>\n<tr>\n<td><strong><code>backoff_base_delay</code></strong></td>\n<td>Optional. <code>int</code> <em>(default: <code>2</code>)</em></td>\n<td>Seconds to wait before the first retry when the provider throttles a request.</td>\n</tr>\n<tr>\n<td><strong><code>backoff_max_attempts</code></strong></td>\n<td>Optional. <code>int</code> <em>(default: <code>3</code>)</em></td>\n<td>Total tries (initial call + retries) before surfacing an error.</td>\n</tr>\n<tr>\n<td><strong><code>backoff_exponential_factor</code></strong></td>\n<td>Optional. <code>int</code> <em>(default: <code>2</code>)</em></td>\n<td>Multiplier that increases the wait time for each retry (<code>delay = base_delay * factor^attempt</code>).</td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"32-example-usage\">3.2 Example Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">llm_config = LLMConfig(\n    provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>,\n    api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>),\n    backoff_base_delay=1, <span class=\"hljs-comment\"># optional</span>\n    backoff_max_attempts=5, <span class=\"hljs-comment\"># optional</span>\n    backoff_exponential_factor=3, <span class=\"hljs-comment\"># optional</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"4-putting-it-all-together\">4. Putting It All Together</h2>\n<ul>\n<li><strong>Use</strong> <code>BrowserConfig</code> for <strong>global</strong> browser settings: engine, headless, proxy, user agent.  </li>\n<li><strong>Use</strong> <code>CrawlerRunConfig</code> for each crawl’s <strong>context</strong>: how to filter content, handle caching, wait for dynamic elements, or run JS.  </li>\n<li><strong>Pass</strong> both configs to <code>AsyncWebCrawler</code> (the <code>BrowserConfig</code>) and then to <code>arun()</code> (the <code>CrawlerRunConfig</code>).  </li>\n<li><strong>Use</strong> <code>LLMConfig</code> for LLM provider configurations that can be used across all extraction, filtering, schema generation, and adaptive crawling tasks. Can be used in - <code>LLMExtractionStrategy</code>, <code>LLMContentFilter</code>, <code>JsonCssExtractionStrategy.generate_schema</code>, <code>JsonXPathExtractionStrategy.generate_schema</code>, and <code>AdaptiveConfig</code> (<code>embedding_llm_config</code> / <code>query_llm_config</code>)</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"># <span class=\"hljs-keyword\">Create</span> a modified <span class=\"hljs-keyword\">copy</span> <span class=\"hljs-keyword\">with</span> the clone() <span class=\"hljs-keyword\">method</span>\nstream_cfg <span class=\"hljs-operator\">=</span> run_cfg.clone(\n    stream<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>,\n    cache_mode<span class=\"hljs-operator\">=</span>CacheMode.BYPASS\n)\n\n# <span class=\"hljs-keyword\">Or</span> <span class=\"hljs-keyword\">set</span> project<span class=\"hljs-operator\">-</span>wide defaults once <span class=\"hljs-keyword\">at</span> startup\nBrowserConfig.set_defaults(headless<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>, text_mode<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>)\nCrawlerRunConfig.set_defaults(cache_mode<span class=\"hljs-operator\">=</span>CacheMode.BYPASS)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/api/strategies/",
    "title": "Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"extraction-chunking-strategies-api\">Extraction &amp; Chunking Strategies API</h1>\n<p>This documentation covers the API reference for extraction and chunking strategies in Crawl4AI.</p>\n<h2 id=\"extraction-strategies\">Extraction Strategies</h2>\n<p>All extraction strategies inherit from the base <code>ExtractionStrategy</code> class and implement two key methods:\n- <code>extract(url: str, html: str) -&gt; List[Dict[str, Any]]</code>\n- <code>run(url: str, sections: List[str]) -&gt; List[Dict[str, Any]]</code></p>\n<h3 id=\"llmextractionstrategy\">LLMExtractionStrategy</h3>\n<p>Used for extracting structured data using Language Models.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">LLMExtractionStrategy(\n    <span class=\"hljs-comment\"># Required Parameters</span>\n    provider: <span class=\"hljs-built_in\">str</span> = DEFAULT_PROVIDER,     <span class=\"hljs-comment\"># LLM provider (e.g., \"ollama/llama2\")</span>\n    api_token: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>,      <span class=\"hljs-comment\"># API token</span>\n\n    <span class=\"hljs-comment\"># Extraction Configuration</span>\n    instruction: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>,              <span class=\"hljs-comment\"># Custom extraction instruction</span>\n    schema: <span class=\"hljs-type\">Dict</span> = <span class=\"hljs-literal\">None</span>,                  <span class=\"hljs-comment\"># Pydantic model schema for structured data</span>\n    extraction_type: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">\"block\"</span>,       <span class=\"hljs-comment\"># \"block\" or \"schema\"</span>\n\n    <span class=\"hljs-comment\"># Chunking Parameters</span>\n    chunk_token_threshold: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">4000</span>,    <span class=\"hljs-comment\"># Maximum tokens per chunk</span>\n    overlap_rate: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.1</span>,           <span class=\"hljs-comment\"># Overlap between chunks</span>\n    word_token_rate: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.75</span>,       <span class=\"hljs-comment\"># Word to token conversion rate</span>\n    apply_chunking: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">True</span>,         <span class=\"hljs-comment\"># Enable/disable chunking</span>\n\n    <span class=\"hljs-comment\"># API Configuration</span>\n    base_url: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>,                <span class=\"hljs-comment\"># Base URL for API</span>\n    extra_args: <span class=\"hljs-type\">Dict</span> = {},               <span class=\"hljs-comment\"># Additional provider arguments</span>\n    verbose: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>                <span class=\"hljs-comment\"># Enable verbose logging</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"regexextractionstrategy\">RegexExtractionStrategy</h3>\n<p>Used for fast pattern-based extraction of common entities using regular expressions.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">RegexExtractionStrategy(\n    <span class=\"hljs-comment\"># Pattern Configuration</span>\n    pattern: IntFlag = RegexExtractionStrategy.Nothing,  <span class=\"hljs-comment\"># Bit flags of built-in patterns to use</span>\n    custom: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-built_in\">str</span>]] = <span class=\"hljs-literal\">None</span>,           <span class=\"hljs-comment\"># Custom pattern dictionary {label: regex}</span>\n\n    <span class=\"hljs-comment\"># Input Format</span>\n    input_format: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">\"fit_html\"</span>,                    <span class=\"hljs-comment\"># \"html\", \"markdown\", \"text\" or \"fit_html\"</span>\n)\n\n<span class=\"hljs-comment\"># Built-in Patterns as Bit Flags</span>\nRegexExtractionStrategy.Email           <span class=\"hljs-comment\"># Email addresses</span>\nRegexExtractionStrategy.PhoneIntl       <span class=\"hljs-comment\"># International phone numbers </span>\nRegexExtractionStrategy.PhoneUS         <span class=\"hljs-comment\"># US-format phone numbers</span>\nRegexExtractionStrategy.Url             <span class=\"hljs-comment\"># HTTP/HTTPS URLs</span>\nRegexExtractionStrategy.IPv4            <span class=\"hljs-comment\"># IPv4 addresses</span>\nRegexExtractionStrategy.IPv6            <span class=\"hljs-comment\"># IPv6 addresses</span>\nRegexExtractionStrategy.Uuid            <span class=\"hljs-comment\"># UUIDs</span>\nRegexExtractionStrategy.Currency        <span class=\"hljs-comment\"># Currency values (USD, EUR, etc)</span>\nRegexExtractionStrategy.Percentage      <span class=\"hljs-comment\"># Percentage values</span>\nRegexExtractionStrategy.Number          <span class=\"hljs-comment\"># Numeric values</span>\nRegexExtractionStrategy.DateIso         <span class=\"hljs-comment\"># ISO format dates</span>\nRegexExtractionStrategy.DateUS          <span class=\"hljs-comment\"># US format dates</span>\nRegexExtractionStrategy.Time24h         <span class=\"hljs-comment\"># 24-hour format times</span>\nRegexExtractionStrategy.PostalUS        <span class=\"hljs-comment\"># US postal codes</span>\nRegexExtractionStrategy.PostalUK        <span class=\"hljs-comment\"># UK postal codes</span>\nRegexExtractionStrategy.HexColor        <span class=\"hljs-comment\"># HTML hex color codes</span>\nRegexExtractionStrategy.TwitterHandle   <span class=\"hljs-comment\"># Twitter handles</span>\nRegexExtractionStrategy.Hashtag         <span class=\"hljs-comment\"># Hashtags</span>\nRegexExtractionStrategy.MacAddr         <span class=\"hljs-comment\"># MAC addresses</span>\nRegexExtractionStrategy.Iban            <span class=\"hljs-comment\"># International bank account numbers</span>\nRegexExtractionStrategy.CreditCard      <span class=\"hljs-comment\"># Credit card numbers</span>\nRegexExtractionStrategy.All             <span class=\"hljs-comment\"># All available patterns</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"cosinestrategy\">CosineStrategy</h3>\n<p>Used for content similarity-based extraction and clustering.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">CosineStrategy(\n    <span class=\"hljs-comment\"># Content Filtering</span>\n    semantic_filter: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>,        <span class=\"hljs-comment\"># Topic/keyword filter</span>\n    word_count_threshold: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">10</span>,     <span class=\"hljs-comment\"># Minimum words per cluster</span>\n    sim_threshold: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.3</span>,         <span class=\"hljs-comment\"># Similarity threshold</span>\n\n    <span class=\"hljs-comment\"># Clustering Parameters</span>\n    max_dist: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.2</span>,             <span class=\"hljs-comment\"># Maximum cluster distance</span>\n    linkage_method: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">'ward'</span>,       <span class=\"hljs-comment\"># Clustering method</span>\n    top_k: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">3</span>,                    <span class=\"hljs-comment\"># Top clusters to return</span>\n\n    <span class=\"hljs-comment\"># Model Configuration</span>\n    model_name: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">'sentence-transformers/all-MiniLM-L6-v2'</span>,  <span class=\"hljs-comment\"># Embedding model</span>\n\n    verbose: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>              <span class=\"hljs-comment\"># Enable verbose logging</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"jsoncssextractionstrategy\">JsonCssExtractionStrategy</h3>\n<p>Used for CSS selector-based structured data extraction.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">JsonCssExtractionStrategy(\n    schema: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">Any</span>],    <span class=\"hljs-comment\"># Extraction schema</span>\n    verbose: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>      <span class=\"hljs-comment\"># Enable verbose logging</span>\n)\n\n<span class=\"hljs-comment\"># Schema Structure</span>\nschema = {\n    <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-built_in\">str</span>,              <span class=\"hljs-comment\"># Schema name</span>\n    <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-built_in\">str</span>,      <span class=\"hljs-comment\"># Base CSS selector</span>\n    <span class=\"hljs-string\">\"fields\"</span>: [               <span class=\"hljs-comment\"># List of fields to extract</span>\n        {\n            <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-built_in\">str</span>,      <span class=\"hljs-comment\"># Field name</span>\n            <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-built_in\">str</span>,  <span class=\"hljs-comment\"># CSS selector</span>\n            <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-built_in\">str</span>,     <span class=\"hljs-comment\"># Field type: \"text\", \"attribute\", \"html\", \"regex\"</span>\n            <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-built_in\">str</span>, <span class=\"hljs-comment\"># For type=\"attribute\"</span>\n            <span class=\"hljs-string\">\"pattern\"</span>: <span class=\"hljs-built_in\">str</span>,  <span class=\"hljs-comment\"># For type=\"regex\"</span>\n            <span class=\"hljs-string\">\"transform\"</span>: <span class=\"hljs-built_in\">str</span>, <span class=\"hljs-comment\"># Optional: \"lowercase\", \"uppercase\", \"strip\"</span>\n            <span class=\"hljs-string\">\"default\"</span>: <span class=\"hljs-type\">Any</span>,   <span class=\"hljs-comment\"># Default value if extraction fails</span>\n            <span class=\"hljs-string\">\"source\"</span>: <span class=\"hljs-built_in\">str</span>,   <span class=\"hljs-comment\"># Optional: navigate to sibling first, e.g. \"+ tr\"</span>\n        }\n    ]\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"chunking-strategies\">Chunking Strategies</h2>\n<p>All chunking strategies inherit from <code>ChunkingStrategy</code> and implement the <code>chunk(text: str) -&gt; list</code> method.</p>\n<h3 id=\"regexchunking\">RegexChunking</h3>\n<p>Splits text based on regex patterns.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">RegexChunking(\n    patterns: <span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>  <span class=\"hljs-comment\"># Regex patterns for splitting</span>\n                               <span class=\"hljs-comment\"># Default: [r'\\n\\n']</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"slidingwindowchunking\">SlidingWindowChunking</h3>\n<p>Creates overlapping chunks with a sliding window approach.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">SlidingWindowChunking(\n    window_size: <span class=\"hljs-keyword\">int</span> = <span class=\"hljs-number\">100</span>,    <span class=\"hljs-comment\"># Window size in words</span>\n    step: <span class=\"hljs-keyword\">int</span> = <span class=\"hljs-number\">50</span>             <span class=\"hljs-comment\"># Step size between windows</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"overlappingwindowchunking\">OverlappingWindowChunking</h3>\n<p>Creates chunks with specified overlap.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-yaml\"><span class=\"hljs-string\">OverlappingWindowChunking(</span>\n    <span class=\"hljs-attr\">window_size:</span> <span class=\"hljs-string\">int</span> <span class=\"hljs-string\">=</span> <span class=\"hljs-number\">1000</span><span class=\"hljs-string\">,</span>   <span class=\"hljs-comment\"># Chunk size in words</span>\n    <span class=\"hljs-attr\">overlap:</span> <span class=\"hljs-string\">int</span> <span class=\"hljs-string\">=</span> <span class=\"hljs-number\">100</span>         <span class=\"hljs-comment\"># Overlap size in words</span>\n<span class=\"hljs-string\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"usage-examples\">Usage Examples</h2>\n<h3 id=\"llm-extraction\">LLM Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Define schema</span>\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Article</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    title: <span class=\"hljs-built_in\">str</span>\n    content: <span class=\"hljs-built_in\">str</span>\n    author: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-comment\"># Create strategy</span>\nstrategy = LLMExtractionStrategy(\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama2\"</span>),\n    schema=Article.schema(),\n    instruction=<span class=\"hljs-string\">\"Extract article details\"</span>\n)\n\n<span class=\"hljs-comment\"># Use with crawler</span>\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com/article\"</span>,\n    extraction_strategy=strategy\n)\n\n<span class=\"hljs-comment\"># Access extracted data</span>\ndata = json.loads(result.extracted_content)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"regex-extraction\">Regex Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, RegexExtractionStrategy\n\n<span class=\"hljs-comment\"># Method 1: Use built-in patterns</span>\nstrategy = RegexExtractionStrategy(\n    pattern = RegexExtractionStrategy.Email | RegexExtractionStrategy.Url\n)\n\n<span class=\"hljs-comment\"># Method 2: Use custom patterns</span>\nprice_pattern = {<span class=\"hljs-string\">\"usd_price\"</span>: <span class=\"hljs-string\">r\"\\$\\s?\\d{1,3}(?:,\\d{3})*(?:\\.\\d{2})?\"</span>}\nstrategy = RegexExtractionStrategy(custom=price_pattern)\n\n<span class=\"hljs-comment\"># Method 3: Generate pattern with LLM assistance (one-time)</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Get sample HTML first</span>\n    sample_result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/products\"</span>)\n    html = sample_result.markdown.fit_html\n\n    <span class=\"hljs-comment\"># Generate regex pattern once</span>\n    pattern = RegexExtractionStrategy.generate_pattern(\n        label=<span class=\"hljs-string\">\"price\"</span>,\n        html=html,\n        query=<span class=\"hljs-string\">\"Product prices in USD format\"</span>,\n        llm_config=LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>)\n    )\n\n    <span class=\"hljs-comment\"># Save pattern for reuse</span>\n    <span class=\"hljs-keyword\">import</span> json\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"price_pattern.json\"</span>, <span class=\"hljs-string\">\"w\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        json.dump(pattern, f)\n\n    <span class=\"hljs-comment\"># Use pattern for extraction (no LLM calls)</span>\n    strategy = RegexExtractionStrategy(custom=pattern)\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n        config=CrawlerRunConfig(extraction_strategy=strategy)\n    )\n\n    <span class=\"hljs-comment\"># Process results</span>\n    data = json.loads(result.extracted_content)\n    <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'label'</span>]}</span>: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"css-extraction\">CSS Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai import JsonCssExtractionStrategy\n\n<span class=\"hljs-comment\"># Define schema</span>\nschema = {\n    <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Product List\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\".product-card\"</span>,\n    <span class=\"hljs-string\">\"fields\"</span>: [\n        {\n            <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2.title\"</span>,\n            <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n        },\n        {\n            <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"price\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".price\"</span>,\n            <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            <span class=\"hljs-string\">\"transform\"</span>: <span class=\"hljs-string\">\"strip\"</span>\n        },\n        {\n            <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"image\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"img\"</span>,\n            <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>,\n            <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"src\"</span>\n        }\n    ]\n}\n\n<span class=\"hljs-comment\"># Create and use strategy</span>\nstrategy = JsonCssExtractionStrategy(schema)\nresult = await crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n    extraction_strategy=strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"content-chunking\">Content Chunking</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.chunking_strategy <span class=\"hljs-keyword\">import</span> OverlappingWindowChunking\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Create chunking strategy</span>\nchunker = OverlappingWindowChunking(\n    window_size=<span class=\"hljs-number\">500</span>,  <span class=\"hljs-comment\"># 500 words per chunk</span>\n    overlap=<span class=\"hljs-number\">50</span>        <span class=\"hljs-comment\"># 50 words overlap</span>\n)\n\n<span class=\"hljs-comment\"># Use with extraction strategy</span>\nstrategy = LLMExtractionStrategy(\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama2\"</span>),\n    chunking_strategy=chunker\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com/long-article\"</span>,\n    extraction_strategy=strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices\">Best Practices</h2>\n<ol>\n<li><strong>Choose the Right Strategy</strong></li>\n<li>Use <code>RegexExtractionStrategy</code> for common data types like emails, phones, URLs, dates</li>\n<li>Use <code>JsonCssExtractionStrategy</code> for well-structured HTML with consistent patterns</li>\n<li>Use <code>LLMExtractionStrategy</code> for complex, unstructured content requiring reasoning</li>\n<li>\n<p>Use <code>CosineStrategy</code> for content similarity and clustering</p>\n</li>\n<li>\n<p><strong>Strategy Selection Guide</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">Is the target data <span class=\"hljs-selector-tag\">a</span> common type (email/phone/date/URL)? \n→ RegexExtractionStrategy\n\nDoes the page have consistent <span class=\"hljs-selector-tag\">HTML</span> structure?\n→ JsonCssExtractionStrategy or JsonXPathExtractionStrategy\n\nIs the data semantically complex or unstructured?\n→ LLMExtractionStrategy\n\nNeed <span class=\"hljs-selector-tag\">to</span> find <span class=\"hljs-attribute\">content</span> similar <span class=\"hljs-selector-tag\">to</span> <span class=\"hljs-selector-tag\">a</span> specific topic?\n→ CosineStrategy\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Optimize Chunking</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># For long documents</span>\nstrategy = LLMExtractionStrategy(\n    chunk_token_threshold=2000,  <span class=\"hljs-comment\"># Smaller chunks</span>\n    overlap_rate=0.1           <span class=\"hljs-comment\"># 10% overlap</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Combine Strategies for Best Performance</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># First pass: Extract structure with CSS</span>\ncss_strategy = JsonCssExtractionStrategy(product_schema)\ncss_result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config=CrawlerRunConfig(extraction_strategy=css_strategy))\nproduct_data = json.loads(css_result.extracted_content)\n\n<span class=\"hljs-comment\"># Second pass: Extract specific fields with regex</span>\ndescriptions = [product[<span class=\"hljs-string\">\"description\"</span>] <span class=\"hljs-keyword\">for</span> product <span class=\"hljs-keyword\">in</span> product_data]\nregex_strategy = RegexExtractionStrategy(\n    pattern=RegexExtractionStrategy.Email | RegexExtractionStrategy.PhoneUS,\n    custom={<span class=\"hljs-string\">\"dimension\"</span>: <span class=\"hljs-string\">r\"\\d+x\\d+x\\d+ (?:cm|in)\"</span>}\n)\n\n<span class=\"hljs-comment\"># Process descriptions with regex</span>\n<span class=\"hljs-keyword\">for</span> text <span class=\"hljs-keyword\">in</span> descriptions:\n    matches = regex_strategy.extract(<span class=\"hljs-string\">\"\"</span>, text)  <span class=\"hljs-comment\"># Direct extraction</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Handle Errors</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">try</span>:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        extraction_strategy=strategy\n    )\n    <span class=\"hljs-keyword\">if</span> result.success:\n        content = json.loads(result.extracted_content)\n<span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extraction failed: <span class=\"hljs-subst\">{e}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Monitor Performance</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">strategy <span class=\"hljs-punctuation\">=</span> CosineStrategy<span class=\"hljs-punctuation\">(</span>\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable logging</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">20</span>,  <span class=\"hljs-comment\"># Filter short content</span>\n    top_k<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5</span>  <span class=\"hljs-comment\"># Limit results</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Cache Generated Patterns</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># For RegexExtractionStrategy pattern generation</span>\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> pathlib <span class=\"hljs-keyword\">import</span> Path\n\ncache_dir = Path(<span class=\"hljs-string\">\"./pattern_cache\"</span>)\ncache_dir.mkdir(exist_ok=<span class=\"hljs-literal\">True</span>)\npattern_file = cache_dir / <span class=\"hljs-string\">\"product_pattern.json\"</span>\n\n<span class=\"hljs-keyword\">if</span> pattern_file.exists():\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(pattern_file) <span class=\"hljs-keyword\">as</span> f:\n        pattern = json.load(f)\n<span class=\"hljs-keyword\">else</span>:\n    <span class=\"hljs-comment\"># Generate once with LLM</span>\n    pattern = RegexExtractionStrategy.generate_pattern(...)\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(pattern_file, <span class=\"hljs-string\">\"w\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        json.dump(pattern, f)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/apps/",
    "title": "Demo Apps - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-interactive-apps\">🚀 Crawl4AI Interactive Apps</h1>\n<p>Welcome to the Crawl4AI Apps Hub - your gateway to interactive tools and demos that make web scraping more intuitive and powerful.</p>\n\n\n<div class=\"intro-section\">\n<h2 id=\"toc-heading-0--interactive-tools-for-modern-web-scraping\">🛠️ Interactive Tools for Modern Web Scraping</h2>\n<p>\nOur apps are designed to make Crawl4AI more accessible and powerful. Whether you're learning browser automation, designing extraction strategies, or building complex scrapers, these tools provide visual, interactive ways to work with Crawl4AI's features.\n</p>\n</div>\n\n<h2 id=\"available-apps\">🎯 Available Apps</h2>\n<div class=\"apps-container\">\n\n<div class=\"app-card\">\n    <span class=\"app-status status-available\">Available</span>\n    <h3 id=\"toc-heading-2--c4a-script-interactive-editor\">🎨 C4A-Script Interactive Editor</h3>\n    <p class=\"app-description\">\n        A visual, block-based programming environment for creating browser automation scripts. Perfect for beginners and experts alike!\n    </p>\n    <ul class=\"app-features\">\n        <li>Drag-and-drop visual programming</li>\n        <li>Real-time JavaScript generation</li>\n        <li>Interactive tutorials</li>\n        <li>Export to C4A-Script or JavaScript</li>\n        <li>Live preview capabilities</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"c4a-script/\" class=\"app-btn\" target=\"_blank\">Launch Editor →</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-available\">Available</span>\n    <h3 id=\"toc-heading-3--llm-context-builder\">🧠 LLM Context Builder</h3>\n    <p class=\"app-description\">\n        Generate optimized context files for your favorite LLM when working with Crawl4AI. Get focused, relevant documentation based on your needs.\n    </p>\n    <ul class=\"app-features\">\n        <li>Modular context generation</li>\n        <li>Memory, reasoning &amp; examples perspectives</li>\n        <li>Component-based selection</li>\n        <li>Vibe coding preset</li>\n        <li>Download custom contexts</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"llmtxt/\" class=\"app-btn\" target=\"_blank\">Launch Builder →</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-coming-soon\">Coming Soon</span>\n    <h3 id=\"toc-heading-4--web-scraping-playground\">🕸️ Web Scraping Playground</h3>\n    <p class=\"app-description\">\n        Test your scraping strategies on real websites with instant feedback. See how different configurations affect your results.\n    </p>\n    <ul class=\"app-features\">\n        <li>Live website testing</li>\n        <li>Side-by-side result comparison</li>\n        <li>Performance metrics</li>\n        <li>Export configurations</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"#\" class=\"app-btn disabled\">Coming Soon</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-available\">Available</span>\n    <h3 id=\"toc-heading-5--crawl4ai-assistant-chrome-extension\">🔍 Crawl4AI Assistant (Chrome Extension)</h3>\n    <p class=\"app-description\">\n        Visual schema builder Chrome extension - click on webpage elements to generate extraction schemas and Python code!\n    </p>\n    <ul class=\"app-features\">\n        <li>Visual element selection</li>\n        <li>Container &amp; field selection modes</li>\n        <li>Smart selector generation</li>\n        <li>Complete Python code generation</li>\n        <li>One-click installation</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"crawl4ai-assistant/\" class=\"app-btn\">Install Extension →</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-coming-soon\">Coming Soon</span>\n    <h3 id=\"toc-heading-6--extraction-lab\">🧪 Extraction Lab</h3>\n    <p class=\"app-description\">\n        Experiment with different extraction strategies and see how they perform on your content. Compare LLM vs CSS vs XPath approaches.\n    </p>\n    <ul class=\"app-features\">\n        <li>Strategy comparison tools</li>\n        <li>Performance benchmarks</li>\n        <li>Cost estimation for LLM strategies</li>\n        <li>Best practice recommendations</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"#\" class=\"app-btn disabled\">Coming Soon</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-coming-soon\">Coming Soon</span>\n    <h3 id=\"toc-heading-7--ai-prompt-designer\">🤖 AI Prompt Designer</h3>\n    <p class=\"app-description\">\n        Craft and test prompts for LLM-based extraction. See how different prompts affect extraction quality and costs.\n    </p>\n    <ul class=\"app-features\">\n        <li>Prompt templates library</li>\n        <li>A/B testing interface</li>\n        <li>Token usage calculator</li>\n        <li>Quality metrics</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"#\" class=\"app-btn disabled\">Coming Soon</a>\n    </div>\n</div>\n\n<div class=\"app-card\">\n    <span class=\"app-status status-coming-soon\">Coming Soon</span>\n    <h3 id=\"toc-heading-8--crawl-monitor\">📊 Crawl Monitor</h3>\n    <p class=\"app-description\">\n        Real-time monitoring dashboard for your crawling operations. Track performance, debug issues, and optimize your scrapers.\n    </p>\n    <ul class=\"app-features\">\n        <li>Real-time crawl statistics</li>\n        <li>Error tracking and debugging</li>\n        <li>Resource usage monitoring</li>\n        <li>Historical analytics</li>\n    </ul>\n    <div class=\"app-action\">\n        <a href=\"#\" class=\"app-btn disabled\">Coming Soon</a>\n    </div>\n</div>\n\n</div>\n\n<h2 id=\"why-use-these-apps\">🚀 Why Use These Apps?</h2>\n<h3 id=\"accelerate-learning\">🎯 <strong>Accelerate Learning</strong></h3>\n<p>Visual tools help you understand Crawl4AI's concepts faster than reading documentation alone.</p>\n<h3 id=\"reduce-development-time\">💡 <strong>Reduce Development Time</strong></h3>\n<p>Generate working code instantly instead of writing everything from scratch.</p>\n<h3 id=\"improve-quality\">🔍 <strong>Improve Quality</strong></h3>\n<p>Test and refine your approach before deploying to production.</p>\n<h3 id=\"community-driven\">🤝 <strong>Community Driven</strong></h3>\n<p>These tools are built based on user feedback. Have an idea? <a href=\"https://github.com/unclecode/crawl4ai/issues\">Let us know</a>!</p>\n<h2 id=\"stay-updated\">📢 Stay Updated</h2>\n<p>Want to know when new apps are released? </p>\n<ul>\n<li>⭐ <a href=\"https://github.com/unclecode/crawl4ai\">Star us on GitHub</a> to get notifications</li>\n<li>🐦 Follow <a href=\"https://twitter.com/unclecode\">@unclecode</a> for announcements</li>\n<li>💬 Join our <a href=\"https://discord.gg/crawl4ai\">Discord community</a> for early access</li>\n</ul>\n<hr>\n<div class=\"admonition tip\">\n<p class=\"admonition-title\">Developer Resources</p>\n<p>Building your own tools with Crawl4AI? Check out our <a href=\"../api/async-webcrawler/\">API Reference</a> and <a href=\"../advanced/advanced-features/\">Integration Guide</a> for comprehensive documentation.</p>\n</div>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/apps/llmtxt/build/",
    "title": "Build - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <p>O<strong>Prompt for AI Coding Assistant: Create an Interactive LLM Context Builder Page</strong></p>\n<p><strong>Objective:</strong></p>\n<p>Your task is to create an interactive HTML webpage with JavaScript functionality that allows users to select and combine different <code>crawl4ai</code> LLM context files into a single downloadable Markdown (<code>.md</code>) file. This tool will empower users to craft tailored context for their AI assistants based on their specific needs.</p>\n<p><strong>Core Functionality:</strong></p>\n<ol>\n<li><strong>Display <code>crawl4ai</code> Components:</strong> The page will list all available <code>crawl4ai</code> documentation components.</li>\n<li><strong>Select Context Types:</strong> For each component, users can select which types of context they want to include:<ul>\n<li>Memory (API facts)</li>\n<li>Reasoning (How-to/why)</li>\n<li>Examples (Code snippets)\n(All should be selected by default for each initially selected component).</li>\n</ul>\n</li>\n<li><strong>Special \"Aggregate\" Contexts:</strong> Include options for special, pre-combined contexts:<ul>\n<li>\"Vibe Coding\" (a curated mix for general AI prompting)</li>\n<li>\"All Library Context\" (a comprehensive aggregation of all memory, reasoning, and examples for the entire library).</li>\n</ul>\n</li>\n<li><strong>Fetch and Concatenate:</strong> When the user clicks a \"Download Combined Context\" button:<ul>\n<li>The JavaScript will fetch the content of all selected Markdown files from the server (from a predefined folder, e.g., <code>/llmtxt/</code>).</li>\n<li>It will concatenate the content of these files into a single string.</li>\n</ul>\n</li>\n<li><strong>Client-Side Download:</strong> The concatenated content will be offered to the user as a download (e.g., <code>custom_crawl4ai_context.md</code>).</li>\n</ol>\n<p><strong>Input/Assumptions:</strong></p>\n<ul>\n<li><strong>Context Files Location:</strong> All individual context Markdown files are located on the server in a publicly accessible folder named <code>llmtxt/</code>.</li>\n<li><strong>File Naming Convention:</strong> Files follow the pattern: <code>crawl4ai_{{component_name}}_[memory|reasoning|examples]_content.llm.md</code>.<ul>\n<li><code>{{component_name}}</code> can contain underscores (e.g., <code>deep_crawling</code>, <code>config_objects</code>).</li>\n<li>The special contexts will have names like <code>crawl4ai_vibe_content.llm.md</code> and <code>crawl4ai_all_content.llm.md</code>.</li>\n</ul>\n</li>\n<li><strong>Component List:</strong> You will be provided with a list of <code>crawl4ai</code> components. For this implementation, use the following list:<ul>\n<li><code>core</code></li>\n<li><code>config_objects</code></li>\n<li><code>deep_crawling</code></li>\n<li><code>deployment</code> (covers Installation &amp; Docker Deployment)</li>\n<li><code>extraction</code> (covers Structured Data Extraction)</li>\n<li><code>markdown</code> (covers Markdown Generation Algorithm)</li>\n<li><code>pdf_processing</code></li>\n<li><em>(No separate \"Vibe Coding\" or \"All Library Context\" in this list, as they are special top-level selections)</em></li>\n</ul>\n</li>\n</ul>\n<p><strong>Detailed UI/UX Requirements:</strong></p>\n<ol>\n<li><strong>Main Page Structure:</strong><ul>\n<li><strong>Header:</strong> \"Crawl4AI Interactive LLM Context Builder\"</li>\n<li><strong>Introduction:</strong> Briefly explain the purpose of the tool (from the <code>USING_LLM_CONTEXTS.md</code> content you helped draft: \"Supercharging Your AI Assistant...\").</li>\n<li><strong>Selection Area:</strong><ul>\n<li><strong>Special Aggregate Contexts (Radio Buttons or Prominent Checkboxes):</strong><ul>\n<li>[ ] \"Vibe Coding Context\" (<code>crawl4ai_vibe_content.llm.md</code>)</li>\n<li>[ ] \"All Library Context (Comprehensive)\" (<code>crawl4ai_all_content.llm.md</code>)</li>\n<li><em>Behavior:</em> Selecting one of these might disable individual component selections (or vice-versa) to avoid redundancy, or simply add them to the list. Consider user experience here. A simple approach is that if an aggregate is selected, it's the <em>only</em> thing downloaded.</li>\n</ul>\n</li>\n<li><strong>Individual Component Selection (Table or List of Checkboxes):</strong><ul>\n<li>A section titled \"Select Individual Components &amp; Context Types:\"</li>\n<li>For each component in the provided list:<ul>\n<li>A master checkbox for the component itself (e.g., <code>[ ] Core Functionality</code>). Selected by default.</li>\n<li>Nested checkboxes (indented or grouped) for context types, enabled only if the parent component is checked:<ul>\n<li><code>[x] Memory (API Facts)</code></li>\n<li><code>[x] Reasoning (How-to/Why)</code></li>\n<li><code>[x] Examples (Code Snippets)</code>\n(These three sub-checkboxes should be selected by default if the parent component is selected).</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n</li>\n<li><strong>Action Button:</strong><ul>\n<li>A button: \"Generate &amp; Download Combined Context\"</li>\n</ul>\n</li>\n<li><strong>Status/Feedback Area:</strong> (Optional, but good UX)<ul>\n<li>Display messages like \"Fetching files...\", \"Combining context...\", \"Download starting...\" or error messages.</li>\n</ul>\n</li>\n</ul>\n</li>\n</ol>\n<p><strong>Final Output:</strong></p>\n<ul>\n<li>A single HTML file (e.g., <code>interactive_context_builder.html</code>).</li>\n<li>Associated JavaScript code (can be inline within <code>&lt;script&gt;</code> tags or in a separate <code>.js</code> file).</li>\n<li>Associated CSS code (can be inline within <code>&lt;style&gt;</code> tags or in a separate <code>.css</code> file).</li>\n</ul>\n<p>This interactive tool will greatly enhance the user experience for <code>crawl4ai</code> developers looking to leverage your specialized LLM contexts. Please ensure the JavaScript is robust and provides good user feedback.</p>\n<hr>\n<p>This prompt should give your AI coding assistant a very clear set of requirements and guidelines for building the interactive context builder. Remember to provide it with the list of components as mentioned in the \"Input/Assumptions\" section.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/apps/llmtxt/why/",
    "title": "Supercharging Your AI Assistant: My Journey to Better LLM Contexts for crawl4ai - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"supercharging-your-ai-assistant-my-journey-to-better-llm-contexts-for-crawl4ai\">Supercharging Your AI Assistant: My Journey to Better LLM Contexts for <code>crawl4ai</code></h1>\n<p>When I started diving deep into using AI coding assistants with my own libraries, particularly <code>crawl4ai</code>, I quickly realized that the common approach to providing context via a simple <code>llm.txt</code> or even a beefed-up <code>README.md</code> just wasn't cutting it. This document explains the problems I encountered and how I've tried to create a more effective system for <code>crawl4ai</code>, allowing you (and your AI assistant) to get precisely the right information.</p>\n<h2 id=\"my-frustration-with-standard-llmtxt-files\">My Frustration with Standard <code>llm.txt</code> Files</h2>\n<p>My experience with generic <code>llm.txt</code> files for complex libraries like <code>crawl4ai</code> revealed several pain points:</p>\n<ol>\n<li>\n<p><strong>Information Overload &amp; Lost Focus:</strong> I found that when I threw a massive, monolithic context file at an LLM, it often struggled. The sheer volume of information seemed to dilute its focus. If I asked a specific question about a niche feature, the LLM might get sidetracked by more prominent but currently irrelevant parts of the library. It felt like trying to find a single sentence in a thousand-page novel – the information was <em>there</em>, but not always accessible or prioritized correctly by the AI.</p>\n</li>\n<li>\n<p><strong>The \"What\" Without the \"How\" or \"Why\":</strong> Most <code>llm.txt</code> files I encountered were essentially API dumps – a list of functions, classes, and parameters. This is the \"what\" of a library. But to truly use a library effectively, especially one as flexible as <code>crawl4ai</code>, you need the \"how\" (idiomatic usage patterns, best practices for common tasks) and the \"why\" (the design rationale behind certain features). Without this, I noticed my AI assistant would often generate syntactically correct but practically inefficient or non-idiomatic code. It was guessing the <em>intent</em> and the <em>best way</em> to use the library, and those guesses weren't always right.</p>\n</li>\n<li>\n<p><strong>No Guidance on \"Thinking\" Like an Expert:</strong> A static list of facts doesn't teach an LLM the <em>art</em> of using the library. It doesn't convey the trade-offs an experienced developer considers, the common pitfalls they've learned to avoid, or the clever ways to combine features to solve complex problems. I wanted my AI assistant to not just recall an API, but to help me <em>reason</em> about the best way to build a solution with <code>crawl4ai</code>.</p>\n</li>\n</ol>\n<h2 id=\"inspiration-selective-inclusion-multi-dimensional-understanding\">Inspiration: Selective Inclusion &amp; Multi-Dimensional Understanding</h2>\n<p>I've always admired how libraries like Lodash or jQuery (in its modular days) allowed developers to pick and choose only the parts they needed, resulting in smaller, more focused bundles. This idea of modularity and selective inclusion resonated deeply with me as I thought about LLM context. Why force-feed an LLM the entire library's details when I'm only working on a specific component or task?</p>\n<p>This led me to develop a new approach for <code>crawl4ai</code>: <strong>multi-dimensional, modular contexts</strong>.</p>\n<p>Instead of one giant <code>llm.txt</code>, I've broken down the <code>crawl4ai</code> documentation into:</p>\n<ol>\n<li><strong>Logical Components:</strong> Context is organized around the major functional areas of the library (e.g., Core, Data Extraction, Deep Crawling, Markdown Generation, etc.). This allows you to select context relevant only to the task at hand.</li>\n<li><strong>Three Dimensions of Context for Each Component:</strong><ul>\n<li><strong><code>_memory.md</code> (Foundational Memory):</strong> This is the \"what.\" It contains the precise, factual information about the component's public API, data structures, configuration objects, parameters, and method signatures. It's the detailed, unambiguous reference.</li>\n<li><strong><code>_reasoning.md</code> (Reasoning &amp; Problem-Solving Framework):</strong> This is the \"how\" and \"why.\" It includes design principles, common task workflows with decision guides, best practices, anti-patterns, illustrative code examples solving real problems, and explanations of trade-offs. It aims to guide the LLM in \"thinking\" like an expert <code>crawl4ai</code> user.</li>\n<li><strong><code>_examples.md</code> (Practical Code Examples):</strong> This is pure \"show-me-the-code.\" It's a collection of runnable snippets demonstrating various ways to use the component's features and configurations, with minimal explanatory text. It’s for quickly seeing different patterns in action.</li>\n</ul>\n</li>\n</ol>\n<p><strong>The Goal:</strong>\nMy aim is to provide you with a flexible system. You can give your AI assistant:\n*   Just the <strong>memory</strong> files for quick API lookups.\n*   The <strong>reasoning</strong> files (perhaps with memory) for help designing solutions.\n*   The <strong>examples</strong> files for seeing practical implementations.\n*   A <strong>combination</strong> of these across one or more components tailored to your specific task.\n*   Or, for broader understanding, special aggregate contexts like the \"Vibe Coding\" context or the \"All Library Context.\"</p>\n<p>By providing these structured, multi-faceted contexts, I hope to significantly improve the quality and relevance of the assistance you get when using AI to code with <code>crawl4ai</code>. The following sections will guide you on how to select and use these different context files.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/basic/installation/",
    "title": "Installation 💻 - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"installation\">Installation 💻</h1>\n<p>Crawl4AI offers flexible installation options to suit various use cases. You can install it as a Python package, use it with Docker, or run it as a local server.</p>\n<h2 id=\"option-1-python-package-installation-recommended\">Option 1: Python Package Installation (Recommended)</h2>\n<p>Crawl4AI is now available on PyPI, making installation easier than ever. Choose the option that best fits your needs:</p>\n<h3 id=\"basic-installation\">Basic Installation</h3>\n<p>For basic web crawling and scraping tasks:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">pip install crawl4ai\nplaywright install <span class=\"hljs-comment\"># Install Playwright dependencies</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"installation-with-pytorch\">Installation with PyTorch</h3>\n<p>For advanced text clustering (includes CosineSimilarity cluster strategy):</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[torch]</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"installation-with-transformers\">Installation with Transformers</h3>\n<p>For text summarization and Hugging Face models:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[transformer]</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"full-installation\">Full Installation</h3>\n<p>For all features:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[all]</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"development-installation\">Development Installation</h3>\n<p>For contributors who plan to modify the source code:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">git <span class=\"hljs-built_in\">clone</span> https://github.com/unclecode/crawl4ai.git\n<span class=\"hljs-built_in\">cd</span> crawl4ai\npip install -e <span class=\"hljs-string\">\".[all]\"</span>\nplaywright install <span class=\"hljs-comment\"># Install Playwright dependencies</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>💡 After installation with \"torch\", \"transformer\", or \"all\" options, it's recommended to run the following CLI command to load the required models:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-download-models\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This is optional but will boost the performance and speed of the crawler. You only need to do this once after installation.</p>\n<h2 id=\"playwright-installation-note-for-ubuntu\">Playwright Installation Note for Ubuntu</h2>\n<p>If you encounter issues with Playwright installation on Ubuntu, you may need to install additional dependencies:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">sudo apt-<span class=\"hljs-keyword\">get</span> install -y \\\n    libwoff1 \\\n    libopus0 \\\n    libwebp7 \\\n    libwebpdemux2 \\\n    libenchant<span class=\"hljs-number\">-2</span><span class=\"hljs-number\">-2</span> \\\n    libgudev<span class=\"hljs-number\">-1.0</span><span class=\"hljs-number\">-0</span> \\\n    libsecret<span class=\"hljs-number\">-1</span><span class=\"hljs-number\">-0</span> \\\n    libhyphen0 \\\n    libgdk-pixbuf2<span class=\"hljs-number\">.0</span><span class=\"hljs-number\">-0</span> \\\n    libegl1 \\\n    libnotify4 \\\n    libxslt1<span class=\"hljs-number\">.1</span> \\\n    libevent<span class=\"hljs-number\">-2.1</span><span class=\"hljs-number\">-7</span> \\\n    libgles2 \\\n    libxcomposite1 \\\n    libatk1<span class=\"hljs-number\">.0</span><span class=\"hljs-number\">-0</span> \\\n    libatk-bridge2<span class=\"hljs-number\">.0</span><span class=\"hljs-number\">-0</span> \\\n    libepoxy0 \\\n    libgtk<span class=\"hljs-number\">-3</span><span class=\"hljs-number\">-0</span> \\\n    libharfbuzz-icu0 \\\n    libgstreamer-gl1<span class=\"hljs-number\">.0</span><span class=\"hljs-number\">-0</span> \\\n    libgstreamer-plugins-bad1<span class=\"hljs-number\">.0</span><span class=\"hljs-number\">-0</span> \\\n    gstreamer1<span class=\"hljs-number\">.0</span>-plugins-good \\\n    gstreamer1<span class=\"hljs-number\">.0</span>-plugins-bad \\\n    libxt6 \\\n    libxaw7 \\\n    xvfb \\\n    fonts-noto-color-emoji \\\n    libfontconfig \\\n    libfreetype6 \\\n    xfonts-cyrillic \\\n    xfonts-scalable \\\n    fonts-liberation \\\n    fonts-ipafont-gothic \\\n    fonts-wqy-zenhei \\\n    fonts-tlwg-loma-otf \\\n    fonts-freefont-ttf\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"option-2-using-docker-coming-soon\">Option 2: Using Docker (Coming Soon)</h2>\n<p>Docker support for Crawl4AI is currently in progress and will be available soon. This will allow you to run Crawl4AI in a containerized environment, ensuring consistency across different systems.</p>\n<h2 id=\"option-3-local-server-installation\">Option 3: Local Server Installation</h2>\n<p>For those who prefer to run Crawl4AI as a local server, instructions will be provided once the Docker implementation is complete.</p>\n<h2 id=\"verifying-your-installation\">Verifying Your Installation</h2>\n<p>After installation, you can verify that Crawl4AI is working correctly by running a simple Python script:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://www.example.com\"</span>)\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">500</span>])  <span class=\"hljs-comment\"># Print first 500 characters</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This script should successfully crawl the example website and print the first 500 characters of the extracted content.</p>\n<h2 id=\"getting-help\">Getting Help</h2>\n<p>If you encounter any issues during installation or usage, please check the <a href=\"https://docs.crawl4ai.com/\">documentation</a> or raise an issue on the <a href=\"https://github.com/unclecode/crawl4ai/issues\">GitHub repository</a>.</p>\n<p>Happy crawling! 🕷️🤖</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/complete-sdk-reference/",
    "title": "Crawl4AI Complete SDK Documentation - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-complete-sdk-documentation\">Crawl4AI Complete SDK Documentation</h1>\n<p><strong>Generated:</strong> 2025-10-19 12:56\n<strong>Format:</strong> Ultra-Dense Reference (Optimized for AI Assistants)\n<strong>Crawl4AI Version:</strong> 0.7.4</p>\n<hr>\n<h2 id=\"navigation\">Navigation</h2>\n<ul>\n<li><a href=\"#installation--setup\">Installation &amp; Setup</a></li>\n<li><a href=\"#quick-start\">Quick Start</a></li>\n<li><a href=\"#core-api\">Core API</a></li>\n<li><a href=\"#configuration\">Configuration</a></li>\n<li><a href=\"#crawling-patterns\">Crawling Patterns</a></li>\n<li><a href=\"#content-processing\">Content Processing</a></li>\n<li><a href=\"#extraction-strategies\">Extraction Strategies</a></li>\n<li><a href=\"#advanced-features\">Advanced Features</a></li>\n</ul>\n<hr>\n<h1 id=\"installation-setup\">Installation &amp; Setup</h1>\n<h1 id=\"installation-setup-2023-edition\">Installation &amp; Setup (2023 Edition)</h1>\n<h2 id=\"1-basic-installation\">1. Basic Installation</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">pip install crawl4ai\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"2-initial-setup-diagnostics\">2. Initial Setup &amp; Diagnostics</h2>\n<h3 id=\"21-run-the-setup-command\">2.1 Run the Setup Command</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- Performs OS-level checks (e.g., missing libs on Linux)\n- Confirms your environment is ready to crawl<p></p>\n<h3 id=\"22-diagnostics\">2.2 Diagnostics</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-doctor\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- Check Python version compatibility\n- Verify Playwright installation\n- Inspect environment variables or library conflicts\nIf any issues arise, follow its suggestions (e.g., installing additional system packages) and re-run <code>crawl4ai-setup</code>.<p></p>\n<h2 id=\"3-verifying-installation-a-simple-crawl-skip-this-step-if-you-already-run-crawl4ai-doctor\">3. Verifying Installation: A Simple Crawl (Skip this step if you already run <code>crawl4ai-doctor</code>)</h2>\n<p>Below is a minimal Python script demonstrating a <strong>basic</strong> crawl. It uses our new <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong> for clarity, though no custom settings are passed in this example:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.example.com\"</span>,\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])  <span class=\"hljs-comment\"># Show the first 300 characters of extracted text</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- A headless browser session loads <code>example.com</code>\n- Crawl4AI returns ~300 characters of markdown.<br>\nIf errors occur, rerun <code>crawl4ai-doctor</code> or manually ensure Playwright is installed correctly.<p></p>\n<h2 id=\"4-advanced-installation-optional\">4. Advanced Installation (Optional)</h2>\n<h3 id=\"41-torch-transformers-or-all\">4.1 Torch, Transformers, or All</h3>\n<ul>\n<li><strong>Text Clustering (Torch)</strong><br>\n  <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[torch]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li><strong>Transformers</strong><br>\n  <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[transformer]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li><strong>All Features</strong><br>\n  <div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[all]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-download-models\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h2 id=\"5-docker-experimental\">5. Docker (Experimental)</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">docker pull unclecode/crawl4ai:basic\ndocker run -p 11235:11235 unclecode/crawl4ai:basic\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nYou can then make POST requests to <code>http://localhost:11235/crawl</code> to perform crawls. <strong>Production usage</strong> is discouraged until our new Docker approach is ready (planned in Jan or Feb 2025).<p></p>\n<h2 id=\"6-local-server-mode-legacy\">6. Local Server Mode (Legacy)</h2>\n<h2 id=\"summary\">Summary</h2>\n<p>1. <strong>Install</strong> with <code>pip install crawl4ai</code> and run <code>crawl4ai-setup</code>.\n2. <strong>Diagnose</strong> with <code>crawl4ai-doctor</code> if you see errors.\n3. <strong>Verify</strong> by crawling <code>example.com</code> with minimal <code>BrowserConfig</code> + <code>CrawlerRunConfig</code>.</p>\n<h1 id=\"quick-start\">Quick Start</h1>\n<h1 id=\"getting-started-with-crawl4ai\">Getting Started with Crawl4AI</h1>\n<ol>\n<li>Run your <strong>first crawl</strong> using minimal configuration.  </li>\n<li>Experiment with a simple <strong>CSS-based extraction</strong> strategy.  </li>\n<li>Crawl a <strong>dynamic</strong> page that loads content via JavaScript.</li>\n</ol>\n<h2 id=\"1-introduction\">1. Introduction</h2>\n<ul>\n<li>An asynchronous crawler, <strong><code>AsyncWebCrawler</code></strong>.  </li>\n<li>Configurable browser and run settings via <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong>.  </li>\n<li>Automatic HTML-to-Markdown conversion via <strong><code>DefaultMarkdownGenerator</code></strong> (supports optional filters).  </li>\n<li>Multiple extraction strategies (LLM-based or “traditional” CSS/XPath-based).</li>\n</ul>\n<h2 id=\"2-your-first-crawl\">2. Your First Crawl</h2>\n<p>Here’s a minimal Python script that creates an <strong><code>AsyncWebCrawler</code></strong>, fetches a webpage, and prints the first 300 characters of its Markdown output:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])  <span class=\"hljs-comment\"># Print first 300 chars</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>AsyncWebCrawler</code></strong> launches a headless browser (Chromium by default).\n- It fetches <code>https://example.com</code>.\n- Crawl4AI automatically converts the HTML into Markdown.<p></p>\n<h2 id=\"3-basic-configuration-light-introduction\">3. Basic Configuration (Light Introduction)</h2>\n<p>1. <strong><code>BrowserConfig</code></strong>: Controls browser behavior (headless or full UI, user agent, JavaScript toggles, etc.).<br>\n2. <strong><code>CrawlerRunConfig</code></strong>: Controls how each crawl runs (caching, extraction, timeouts, hooking, etc.).\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_conf = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)  <span class=\"hljs-comment\"># or False to see the browser</span>\n    run_conf = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_conf\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<blockquote>\n<p>IMPORTANT: By default cache mode is set to <code>CacheMode.BYPASS</code> to have fresh content. Set <code>CacheMode.ENABLED</code> to enable caching.</p>\n</blockquote>\n<h2 id=\"4-generating-markdown-output\">4. Generating Markdown Output</h2>\n<ul>\n<li><strong><code>result.markdown</code></strong>:  </li>\n<li><strong><code>result.markdown.fit_markdown</code></strong>:<br>\n  The same content after applying any configured <strong>content filter</strong> (e.g., <code>PruningContentFilter</code>).</li>\n</ul>\n<h3 id=\"example-using-a-filter-with-defaultmarkdowngenerator\">Example: Using a Filter with <code>DefaultMarkdownGenerator</code></h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\nmd_generator = DefaultMarkdownGenerator(\n    content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.4</span>, threshold_type=<span class=\"hljs-string\">\"fixed\"</span>)\n)\n\nconfig = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS,\n    markdown_generator=md_generator\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.fit_markdown))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Note</strong>: If you do <strong>not</strong> specify a content filter or markdown generator, you’ll typically see only the raw Markdown. <code>PruningContentFilter</code> may adds around <code>50ms</code> in processing time. We’ll dive deeper into these strategies in a dedicated <strong>Markdown Generation</strong> tutorial.<p></p>\n<h2 id=\"5-simple-data-extraction-css-based\">5. Simple Data Extraction (CSS-based)</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Generate a schema (one-time cost)</span>\nhtml = <span class=\"hljs-string\">\"&lt;div class='product'&gt;&lt;h2&gt;Gaming Laptop&lt;/h2&gt;&lt;span class='price'&gt;$999.99&lt;/span&gt;&lt;/div&gt;\"</span>\n\n<span class=\"hljs-comment\"># Using OpenAI (requires API token)</span>\nschema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-openai-token\"</span>)  <span class=\"hljs-comment\"># Required for OpenAI</span>\n)\n\n<span class=\"hljs-comment\"># Or using Ollama (open source, no token needed)</span>\nschema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama3.3\"</span>, api_token=<span class=\"hljs-literal\">None</span>)  <span class=\"hljs-comment\"># Not needed for Ollama</span>\n)\n\n<span class=\"hljs-comment\"># Use the schema for fast, repeated extractions</span>\nstrategy = JsonCssExtractionStrategy(schema)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Example Items\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.item\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n\n    raw_html = <span class=\"hljs-string\">\"&lt;div class='item'&gt;&lt;h2&gt;Item 1&lt;/h2&gt;&lt;a href='https://example.com/item1'&gt;Link 1&lt;/a&gt;&lt;/div&gt;\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"raw://\"</span> + raw_html,\n            config=CrawlerRunConfig(\n                cache_mode=CacheMode.BYPASS,\n                extraction_strategy=JsonCssExtractionStrategy(schema)\n            )\n        )\n        <span class=\"hljs-comment\"># The JSON output is stored in 'extracted_content'</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(data)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- Great for repetitive page structures (e.g., item listings, articles).\n- No AI usage or costs.\n- The crawler returns a JSON string you can parse or store.\n- For sites where data is split across sibling elements (e.g. Hacker News), use the <code>\"source\"</code> field key to navigate to a sibling before extracting: <code>{\"name\": \"score\", \"selector\": \"span.score\", \"type\": \"text\", \"source\": \"+ tr\"}</code>.<p></p>\n<blockquote>\n<p>Tips: You can pass raw HTML to the crawler instead of a URL. To do so, prefix the HTML with <code>raw://</code>.</p>\n</blockquote>\n<h2 id=\"6-simple-data-extraction-llm-based\">6. Simple Data Extraction (LLM-based)</h2>\n<ul>\n<li><strong>Open-Source Models</strong> (e.g., <code>ollama/llama3.3</code>, <code>no_token</code>)  </li>\n<li><strong>OpenAI Models</strong> (e.g., <code>openai/gpt-4</code>, requires <code>api_token</code>)  </li>\n<li>Or any provider supported by the underlying library\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">OpenAIModelFee</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    model_name: <span class=\"hljs-built_in\">str</span> = Field(..., description=<span class=\"hljs-string\">\"Name of the OpenAI model.\"</span>)\n    input_fee: <span class=\"hljs-built_in\">str</span> = Field(..., description=<span class=\"hljs-string\">\"Fee for input token for the OpenAI model.\"</span>)\n    output_fee: <span class=\"hljs-built_in\">str</span> = Field(\n        ..., description=<span class=\"hljs-string\">\"Fee for output token for the OpenAI model.\"</span>\n    )\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_structured_data_using_llm</span>(<span class=\"hljs-params\">\n    provider: <span class=\"hljs-built_in\">str</span>, api_token: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>, extra_headers: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n</span>):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n--- Extracting Structured Data with <span class=\"hljs-subst\">{provider}</span> ---\"</span>)\n\n    <span class=\"hljs-keyword\">if</span> api_token <span class=\"hljs-keyword\">is</span> <span class=\"hljs-literal\">None</span> <span class=\"hljs-keyword\">and</span> provider != <span class=\"hljs-string\">\"ollama\"</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"API token is required for <span class=\"hljs-subst\">{provider}</span>. Skipping this example.\"</span>)\n        <span class=\"hljs-keyword\">return</span>\n\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)\n\n    extra_args = {<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0</span>, <span class=\"hljs-string\">\"top_p\"</span>: <span class=\"hljs-number\">0.9</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">2000</span>}\n    <span class=\"hljs-keyword\">if</span> extra_headers:\n        extra_args[<span class=\"hljs-string\">\"extra_headers\"</span>] = extra_headers\n\n    crawler_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        word_count_threshold=<span class=\"hljs-number\">1</span>,\n        page_timeout=<span class=\"hljs-number\">80000</span>,\n        extraction_strategy=LLMExtractionStrategy(\n            llm_config = LLMConfig(provider=provider,api_token=api_token),\n            schema=OpenAIModelFee.model_json_schema(),\n            extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n            instruction=<span class=\"hljs-string\">\"\"\"From the crawled content, extract all mentioned model names along with their fees for input and output tokens. \n            Do not miss any models in the entire content.\"\"\"</span>,\n            extra_args=extra_args,\n        ),\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://openai.com/api/pricing/\"</span>, config=crawler_config\n        )\n        <span class=\"hljs-built_in\">print</span>(result.extracted_content)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n\n    asyncio.run(\n        extract_structured_data_using_llm(\n            provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>, api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>)\n        )\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li>We define a Pydantic schema (<code>PricingInfo</code>) describing the fields we want.</li>\n</ul>\n<h2 id=\"7-adaptive-crawling-new\">7. Adaptive Crawling (New!)</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, AdaptiveCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">adaptive_example</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        adaptive = AdaptiveCrawler(crawler)\n\n        <span class=\"hljs-comment\"># Start adaptive crawling</span>\n        result = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n            start_url=<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>,\n            query=<span class=\"hljs-string\">\"async context managers\"</span>\n        )\n\n        <span class=\"hljs-comment\"># View results</span>\n        adaptive.print_stats()\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.crawled_urls)}</span> pages\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Achieved <span class=\"hljs-subst\">{adaptive.confidence:<span class=\"hljs-number\">.0</span>%}</span> confidence\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(adaptive_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong>Automatic stopping</strong>: Stops when sufficient information is gathered\n- <strong>Intelligent link selection</strong>: Follows only relevant links\n- <strong>Confidence scoring</strong>: Know how complete your information is<p></p>\n<h2 id=\"8-multi-url-concurrency-preview\">8. Multi-URL Concurrency (Preview)</h2>\n<p>If you need to crawl multiple URLs in <strong>parallel</strong>, you can use <code>arun_many()</code>. By default, Crawl4AI employs a <strong>MemoryAdaptiveDispatcher</strong>, automatically adjusting concurrency based on system resources. Here’s a quick glimpse:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">quick_parallel_example</span>():\n    urls = [\n        <span class=\"hljs-string\">\"https://example.com/page1\"</span>,\n        <span class=\"hljs-string\">\"https://example.com/page2\"</span>,\n        <span class=\"hljs-string\">\"https://example.com/page3\"</span>\n    ]\n\n    run_conf = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        stream=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Stream results as they complete</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls, config=run_conf):\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] <span class=\"hljs-subst\">{result.url}</span>, length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown)}</span>\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[ERROR] <span class=\"hljs-subst\">{result.url}</span> =&gt; <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n        <span class=\"hljs-comment\"># Or get all results at once (default behavior)</span>\n        run_conf = run_conf.clone(stream=<span class=\"hljs-literal\">False</span>)\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls, config=run_conf)\n        <span class=\"hljs-keyword\">for</span> res <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-keyword\">if</span> res.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] <span class=\"hljs-subst\">{res.url}</span>, length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(res.markdown.raw_markdown)}</span>\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[ERROR] <span class=\"hljs-subst\">{res.url}</span> =&gt; <span class=\"hljs-subst\">{res.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(quick_parallel_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. <strong>Streaming mode</strong> (<code>stream=True</code>): Process results as they become available using <code>async for</code>\n2. <strong>Batch mode</strong> (<code>stream=False</code>): Wait for all results to complete<p></p>\n<h2 id=\"8-dynamic-content-example\">8. Dynamic Content Example</h2>\n<p>Some sites require multiple “page clicks” or dynamic JavaScript updates. Below is an example showing how to <strong>click</strong> a “Next Page” button and wait for new commits to load on GitHub, using <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_structured_data_using_css_extractor</span>():\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n--- Using JsonCssExtractionStrategy for Fast Structured Output ---\"</span>)\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"KidoCode Courses\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"section.charge-methodology .w-tab-content &gt; div\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"section_title\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h3.heading-50\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"section_description\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".charge-content\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".text-block-93\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_description\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".course-content-text\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_icon\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".image-92\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>,\n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"src\"</span>,\n            },\n        ],\n    }\n\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, java_script_enabled=<span class=\"hljs-literal\">True</span>)\n\n    js_click_tabs = <span class=\"hljs-string\">\"\"\"\n    (async () =&gt; {\n        const tabs = document.querySelectorAll(\"section.charge-methodology .tabs-menu-3 &gt; div\");\n        for(let tab of tabs) {\n            tab.scrollIntoView();\n            tab.click();\n            await new Promise(r =&gt; setTimeout(r, 500));\n        }\n    })();\n    \"\"\"</span>\n\n    crawler_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n        js_code=[js_click_tabs],\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.kidocode.com/degrees/technology\"</span>, config=crawler_config\n        )\n\n        companies = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(companies)}</span> companies\"</span>)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(companies[<span class=\"hljs-number\">0</span>], indent=<span class=\"hljs-number\">2</span>))\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">await</span> extract_structured_data_using_css_extractor()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>BrowserConfig(headless=False)</code></strong>: We want to watch it click “Next Page.”<br>\n- <strong><code>CrawlerRunConfig(...)</code></strong>: We specify the extraction strategy, pass <code>session_id</code> to reuse the same page.<br>\n- <strong><code>js_code</code></strong> and <strong><code>wait_for</code></strong> are used for subsequent pages (<code>page &gt; 0</code>) to click the “Next” button and wait for new commits to load.<br>\n- <strong><code>js_only=True</code></strong> indicates we’re not re-navigating but continuing the existing session.<br>\n- Finally, we call <code>kill_session()</code> to clean up the page and browser session.<p></p>\n<h2 id=\"9-next-steps\">9. Next Steps</h2>\n<ol>\n<li>Performed a basic crawl and printed Markdown.  </li>\n<li>Used <strong>content filters</strong> with a markdown generator.  </li>\n<li>Extracted JSON via <strong>CSS</strong> or <strong>LLM</strong> strategies.  </li>\n<li>Handled <strong>dynamic</strong> pages with JavaScript triggers.</li>\n</ol>\n<h1 id=\"core-api\">Core API</h1>\n<h1 id=\"asyncwebcrawler\">AsyncWebCrawler</h1>\n<p>The <strong><code>AsyncWebCrawler</code></strong> is the core class for asynchronous web crawling in Crawl4AI. You typically create it <strong>once</strong>, optionally customize it with a <strong><code>BrowserConfig</code></strong> (e.g., headless, user agent), then <strong>run</strong> multiple <strong><code>arun()</code></strong> calls with different <strong><code>CrawlerRunConfig</code></strong> objects.\n1. <strong>Create</strong> a <code>BrowserConfig</code> for global browser settings.  \n2. <strong>Instantiate</strong> <code>AsyncWebCrawler(config=browser_config)</code>.  \n3. <strong>Use</strong> the crawler in an async context manager (<code>async with</code>) or manage start/close manually.  \n4. <strong>Call</strong> <code>arun(url, config=crawler_run_config)</code> for each page you want.</p>\n<h2 id=\"1-constructor-overview\">1. Constructor Overview</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">AsyncWebCrawler</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        self,\n        crawler_strategy: <span class=\"hljs-type\">Optional</span>[AsyncCrawlerStrategy] = <span class=\"hljs-literal\">None</span>,\n        config: <span class=\"hljs-type\">Optional</span>[BrowserConfig] = <span class=\"hljs-literal\">None</span>,\n        always_bypass_cache: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>,           <span class=\"hljs-comment\"># deprecated</span>\n        always_by_pass_cache: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">bool</span>] = <span class=\"hljs-literal\">None</span>, <span class=\"hljs-comment\"># also deprecated</span>\n        base_directory: <span class=\"hljs-built_in\">str</span> = ...,\n        thread_safe: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>,\n        **kwargs,\n    </span>):\n        <span class=\"hljs-string\">\"\"\"\n        Create an AsyncWebCrawler instance.\n\n        Args:\n            crawler_strategy: \n                (Advanced) Provide a custom crawler strategy if needed.\n            config: \n                A BrowserConfig object specifying how the browser is set up.\n            always_bypass_cache: \n                (Deprecated) Use CrawlerRunConfig.cache_mode instead.\n            base_directory:     \n                Folder for storing caches/logs (if relevant).\n            thread_safe: \n                If True, attempts some concurrency safeguards. Usually False.\n            **kwargs: \n                Additional legacy or debugging parameters.\n        \"\"\"</span>\n    )\n\n<span class=\"hljs-comment\">### Typical Initialization</span>\n\n```python\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\nbrowser_cfg = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n    headless=<span class=\"hljs-literal\">True</span>,\n    verbose=<span class=\"hljs-literal\">True</span>\ncrawler = AsyncWebCrawler(config=browser_cfg)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Notes</strong>:</p>\n<ul>\n<li><strong>Legacy</strong> parameters like <code>always_bypass_cache</code> remain for backward compatibility, but prefer to set <strong>caching</strong> in <code>CrawlerRunConfig</code>.</li>\n</ul>\n<hr>\n<h2 id=\"2-lifecycle-startclose-or-context-manager\">2. Lifecycle: Start/Close or Context Manager</h2>\n<h3 id=\"21-context-manager-recommended\">2.1 Context Manager (Recommended)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>(<span class=\"hljs-params\">config=browser_cfg</span>) <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n    <span class=\"hljs-meta\"># The crawler automatically starts/closes resources</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>When the <code>async with</code> block ends, the crawler cleans up (closes the browser, etc.).</p>\n<h3 id=\"22-manual-start-close\">2.2 Manual Start &amp; Close</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">crawler = AsyncWebCrawler(config=browser_cfg)\n<span class=\"hljs-keyword\">await</span> crawler.start()\nresult1 = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\nresult2 = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://another.com\"</span>)\n<span class=\"hljs-keyword\">await</span> crawler.close()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Use this style if you have a <strong>long-running</strong> application or need full control of the crawler’s lifecycle.</p>\n<hr>\n<h2 id=\"3-primary-method-arun\">3. Primary Method: <code>arun()</code></h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun</span>(<span class=\"hljs-params\">\n    url: <span class=\"hljs-built_in\">str</span>,\n    config: <span class=\"hljs-type\">Optional</span>[CrawlerRunConfig] = <span class=\"hljs-literal\">None</span>,\n    <span class=\"hljs-comment\"># Legacy parameters for backward compatibility...</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"31-new-approach\">3.1 New Approach</h3>\n<p>You pass a <code>CrawlerRunConfig</code> object that sets up everything about a crawl—content filtering, caching, session reuse, JS code, screenshots, etc.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, CacheMode\nrun_cfg = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS,\n    css_selector=<span class=\"hljs-string\">\"main.article\"</span>,\n    word_count_threshold=<span class=\"hljs-number\">10</span>,\n    screenshot=<span class=\"hljs-literal\">True</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/news\"</span>, config=run_cfg)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-legacy-parameters-still-accepted\">3.2 Legacy Parameters Still Accepted</h3>\n<p>For <strong>backward</strong> compatibility, <code>arun()</code> can still accept direct arguments like <code>css_selector=...</code>, <code>word_count_threshold=...</code>, etc., but we strongly advise migrating them into a <strong><code>CrawlerRunConfig</code></strong>.</p>\n<hr>\n<h2 id=\"4-batch-processing-arun_many\">4. Batch Processing: <code>arun_many()</code></h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun_many</span>(<span class=\"hljs-params\">\n    urls: <span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>],\n    config: <span class=\"hljs-type\">Optional</span>[CrawlerRunConfig] = <span class=\"hljs-literal\">None</span>,\n    <span class=\"hljs-comment\"># Legacy parameters maintained for backwards compatibility...</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"41-resource-aware-crawling\">4.1 Resource-Aware Crawling</h3>\n<p>The <code>arun_many()</code> method now uses an intelligent dispatcher that:</p>\n<ul>\n<li>Monitors system memory usage</li>\n<li>Implements adaptive rate limiting</li>\n<li>Provides detailed progress monitoring</li>\n<li>Manages concurrent crawls efficiently</li>\n</ul>\n<h3 id=\"42-example-usage\">4.2 Example Usage</h3>\n<p>Check page <a href=\"../advanced/multi-url-crawling.md\">Multi-url Crawling</a> for a detailed example of how to use <code>arun_many()</code>.</p>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\">### 4.3 Key Features</span>\n<span class=\"hljs-number\">1.</span> **Rate Limiting**\n   - Automatic delay between requests\n   - Exponential backoff on rate limit detection\n   - Domain-specific rate limiting\n   - Configurable retry strategy\n<span class=\"hljs-number\">2.</span> **Resource Monitoring**\n   - Memory usage tracking\n   - Adaptive concurrency based on system load\n   - Automatic pausing when resources are constrained\n<span class=\"hljs-number\">3.</span> **Progress Monitoring**\n   - Detailed <span class=\"hljs-keyword\">or</span> aggregated progress display\n   - Real-time status updates\n   - Memory usage statistics\n<span class=\"hljs-number\">4.</span> **Error Handling**\n   - Graceful handling of rate limits\n   - Automatic retries <span class=\"hljs-keyword\">with</span> backoff\n   - Detailed error reporting\n<span class=\"hljs-comment\">## 5. `CrawlResult` Output</span>\nEach `arun()` returns a **`CrawlResult`** containing:\n- `url`: Final URL (<span class=\"hljs-keyword\">if</span> redirected).\n- `html`: Original HTML.\n- `cleaned_html`: Sanitized HTML.\n- `markdown_v2`: Removed <span class=\"hljs-keyword\">in</span> v0<span class=\"hljs-number\">.5</span>. Accessing it raises `AttributeError`. Use `markdown`.\n- `extracted_content`: If an extraction strategy was used (JSON <span class=\"hljs-keyword\">for</span> CSS/LLM strategies).\n- `screenshot`, `pdf`: If screenshots/PDF requested.\n- `media`, `links`: Information about discovered images/links.\n- `success`, `error_message`: Status info.\n<span class=\"hljs-comment\">## 6. Quick Example</span>\n```python\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">import</span> json\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Browser config</span>\n    browser_cfg = BrowserConfig(\n        browser_type=<span class=\"hljs-string\">\"firefox\"</span>,\n        headless=<span class=\"hljs-literal\">False</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># 2. Run config</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"article.post\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"url\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, \n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>\n            }\n        ]\n    }\n\n    run_cfg = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n        word_count_threshold=<span class=\"hljs-number\">15</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.post\"</span>  <span class=\"hljs-comment\"># Wait for posts to appear</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/blog\"</span>,\n            config=run_cfg\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-keyword\">if</span> result.extracted_content:\n                articles = json.loads(result.extracted_content)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted articles:\"</span>, articles[:<span class=\"hljs-number\">2</span>])\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- We define a <strong><code>BrowserConfig</code></strong> with Firefox, no headless, and <code>verbose=True</code>.  \n- We define a <strong><code>CrawlerRunConfig</code></strong> that <strong>bypasses cache</strong>, uses a <strong>CSS</strong> extraction schema, has a <code>word_count_threshold=15</code>, etc.  \n- We pass them to <code>AsyncWebCrawler(config=...)</code> and <code>arun(url=..., config=...)</code>.<p></p>\n<h2 id=\"7-best-practices-migration-notes\">7. Best Practices &amp; Migration Notes</h2>\n<p>1. <strong>Use</strong> <code>BrowserConfig</code> for <strong>global</strong> settings about the browser’s environment.  \n2. <strong>Use</strong> <code>CrawlerRunConfig</code> for <strong>per-crawl</strong> logic (caching, content filtering, extraction strategies, wait conditions).  \n3. <strong>Avoid</strong> legacy parameters like <code>css_selector</code> or <code>word_count_threshold</code> directly in <code>arun()</code>. Instead:\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">run_cfg</span> = CrawlerRunConfig(css_selector=<span class=\"hljs-string\">\".main-content\"</span>, word_count_threshold=<span class=\"hljs-number\">20</span>)\n<span class=\"hljs-attr\">result</span> = await crawler.arun(url=<span class=\"hljs-string\">\"...\"</span>, config=run_cfg)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"8-summary\">8. Summary</h2>\n<ul>\n<li><strong>Constructor</strong> accepts <strong><code>BrowserConfig</code></strong> (or defaults).  </li>\n<li><strong><code>arun(url, config=CrawlerRunConfig)</code></strong> is the main method for single-page crawls.  </li>\n<li><strong><code>arun_many(urls, config=CrawlerRunConfig)</code></strong> handles concurrency across multiple URLs.  </li>\n<li>For advanced lifecycle control, use <code>start()</code> and <code>close()</code> explicitly.  </li>\n<li>If you used <code>AsyncWebCrawler(browser_type=\"chromium\", css_selector=\"...\")</code>, move browser settings to <code>BrowserConfig(...)</code> and content/crawl logic to <code>CrawlerRunConfig(...)</code>.</li>\n</ul>\n<h1 id=\"arun-parameter-guide-new-approach\"><code>arun()</code> Parameter Guide (New Approach)</h1>\n<p>In Crawl4AI’s <strong>latest</strong> configuration model, nearly all parameters that once went directly to <code>arun()</code> are now part of <strong><code>CrawlerRunConfig</code></strong>. When calling <code>arun()</code>, you provide:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,  \n    config=my_run_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nBelow is an organized look at the parameters that can go inside <code>CrawlerRunConfig</code>, divided by their functional areas. For <strong>Browser</strong> settings (e.g., <code>headless</code>, <code>browser_type</code>), see <a href=\"./parameters.md\">BrowserConfig</a>.<p></p>\n<h2 id=\"1-core-usage\">1. Core Usage</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    run_config = CrawlerRunConfig(\n        verbose=<span class=\"hljs-literal\">True</span>,            <span class=\"hljs-comment\"># Detailed logging</span>\n        cache_mode=CacheMode.ENABLED,  <span class=\"hljs-comment\"># Use normal read/write cache</span>\n        check_robots_txt=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Respect robots.txt rules</span>\n        <span class=\"hljs-comment\"># ... other parameters</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n\n        <span class=\"hljs-comment\"># Check if blocked by robots.txt</span>\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success <span class=\"hljs-keyword\">and</span> result.status_code == <span class=\"hljs-number\">403</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>verbose=True</code> logs each crawl step.  \n- <code>cache_mode</code> decides how to read/write the local crawl cache.<p></p>\n<h2 id=\"2-cache-control\">2. Cache Control</h2>\n<p><strong><code>cache_mode</code></strong> (default: <code>CacheMode.ENABLED</code>)<br>\nUse a built-in enum from <code>CacheMode</code>:\n- <code>ENABLED</code>: Normal caching—reads if available, writes if missing.\n- <code>DISABLED</code>: No caching—always refetch pages.\n- <code>READ_ONLY</code>: Reads from cache only; no new writes.\n- <code>WRITE_ONLY</code>: Writes to cache but doesn’t read existing data.\n- <code>BYPASS</code>: Skips reading cache for this crawl (though it might still write if set up that way).\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>CacheMode.BYPASS</code> — Skip cache entirely; always fetch fresh, write result to cache.\n- <code>CacheMode.DISABLED</code> — No caching at all; don't read or write.\n- <code>CacheMode.WRITE_ONLY</code> — Never read from cache, but write results.\n- <code>CacheMode.READ_ONLY</code> — Read from cache if available, never write.<p></p>\n<h2 id=\"3-content-processing-selection\">3. Content Processing &amp; Selection</h2>\n<h3 id=\"31-text-processing\">3.1 Text Processing</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">10</span>,   <span class=\"hljs-comment\"># Ignore text blocks &lt;10 words</span>\n    only_text<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>,           <span class=\"hljs-comment\"># If True, tries to remove non-text elements</span>\n    keep_data_attributes<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span> <span class=\"hljs-comment\"># Keep or discard data-* attributes</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-content-selection\">3.2 Content Selection</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    css_selector<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\".main-content\"</span>,  <span class=\"hljs-comment\"># Focus on .main-content region only</span>\n    excluded_tags<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"form\"</span>, <span class=\"hljs-string\">\"nav\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Remove entire tag blocks</span>\n    remove_forms<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># Specifically strip &lt;form&gt; elements</span>\n    remove_overlay_elements<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Attempt to remove modals/popups</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"33-link-handling\">3.3 Link Handling</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,         <span class=\"hljs-comment\"># Remove external links from final content</span>\n    exclude_social_media_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,     <span class=\"hljs-comment\"># Remove links to known social sites</span>\n    exclude_domains<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"ads.example.com\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Exclude links to these domains</span>\n    exclude_social_media_domains<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"facebook.com\"</span>,<span class=\"hljs-string\">\"twitter.com\"</span><span class=\"hljs-punctuation\">]</span>, <span class=\"hljs-comment\"># Extend the default list</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"34-media-filtering\">3.4 Media Filtering</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Strip images from other domains</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"4-page-navigation-timing\">4. Page Navigation &amp; Timing</h2>\n<h3 id=\"41-basic-browser-flow\">4.1 Basic Browser Flow</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    wait_for=<span class=\"hljs-string\">\"css:.dynamic-content\"</span>, <span class=\"hljs-comment\"># Wait for .dynamic-content</span>\n    delay_before_return_html=2.0,    <span class=\"hljs-comment\"># Wait 2s before capturing final HTML</span>\n    page_timeout=60000,             <span class=\"hljs-comment\"># Navigation &amp; script timeout (ms)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>wait_for</code>:<br>\n  - <code>\"css:selector\"</code> or<br>\n  - <code>\"js:() =&gt; boolean\"</code><br>\n  e.g. <code>js:() =&gt; document.querySelectorAll('.item').length &gt; 10</code>.\n- <code>mean_delay</code> &amp; <code>max_range</code>: define random delays for <code>arun_many()</code> calls.  \n- <code>semaphore_count</code>: concurrency limit when crawling multiple URLs.<p></p>\n<h3 id=\"42-javascript-execution\">4.2 JavaScript Execution</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    js_code<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        <span class=\"hljs-string\">\"document.querySelector('.load-more')?.click();\"</span>\n    <span class=\"hljs-punctuation\">]</span>,\n    js_only<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>js_code</code> can be a single string or a list of strings.  \n- <code>js_only=True</code> means “I’m continuing in the same session with new JS steps, no new full navigation.”<p></p>\n<h3 id=\"43-anti-bot\">4.3 Anti-Bot</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    magic<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    simulate_user<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    override_navigator<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>magic=True</code> tries multiple stealth features.  \n- <code>simulate_user=True</code> mimics mouse movements or random delays.  \n- <code>override_navigator=True</code> fakes some navigator properties (like user agent checks).<p></p>\n<h2 id=\"5-session-management\">5. Session Management</h2>\n<p><strong><code>session_id</code></strong>: \n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    session_id=<span class=\"hljs-string\">\"my_session123\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nIf re-used in subsequent <code>arun()</code> calls, the same tab/page context is continued (helpful for multi-step tasks or stateful browsing).<p></p>\n<h2 id=\"6-screenshot-pdf-media-options\">6. Screenshot, PDF &amp; Media Options</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    screenshot<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># Grab a screenshot as base64</span>\n    screenshot_wait_for<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">1.0</span>,     <span class=\"hljs-comment\"># Wait 1s before capturing</span>\n    pdf<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Also produce a PDF</span>\n    image_description_min_word_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5</span>,  <span class=\"hljs-comment\"># If analyzing alt text</span>\n    image_score_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3</span>,                <span class=\"hljs-comment\"># Filter out low-score images</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>result.screenshot</code> → Base64 screenshot string.\n- <code>result.pdf</code> → Byte array with PDF data.<p></p>\n<h2 id=\"7-extraction-strategy\">7. Extraction Strategy</h2>\n<p><strong>For advanced data extraction</strong> (CSS/LLM-based), set <code>extraction_strategy</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">run_config = CrawlerRunConfig(\n    extraction_strategy=my_css_or_llm_strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThe extracted data will appear in <code>result.extracted_content</code>.<p></p>\n<h2 id=\"8-comprehensive-example\">8. Comprehensive Example</h2>\n<p>Below is a snippet combining many parameters:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Example schema</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"article.post\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>,  <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>,  <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n\n    run_config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Core</span>\n        verbose=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.ENABLED,\n        check_robots_txt=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Respect robots.txt rules</span>\n\n        <span class=\"hljs-comment\"># Content</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        css_selector=<span class=\"hljs-string\">\"main.content\"</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Page &amp; JS</span>\n        js_code=<span class=\"hljs-string\">\"document.querySelector('.show-more')?.click();\"</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.loaded-block\"</span>,\n        page_timeout=<span class=\"hljs-number\">30000</span>,\n\n        <span class=\"hljs-comment\"># Extraction</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n\n        <span class=\"hljs-comment\"># Session</span>\n        session_id=<span class=\"hljs-string\">\"persistent_session\"</span>,\n\n        <span class=\"hljs-comment\"># Media</span>\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        pdf=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Anti-bot</span>\n        simulate_user=<span class=\"hljs-literal\">True</span>,\n        magic=<span class=\"hljs-literal\">True</span>,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/posts\"</span>, config=run_config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extraction JSON:\"</span>, result.extracted_content)\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n            <span class=\"hljs-keyword\">if</span> result.pdf:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"PDF bytes length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.pdf))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. <strong>Crawling</strong> the main content region, ignoring external links.  \n2. Running <strong>JavaScript</strong> to click “.show-more”.  \n3. <strong>Waiting</strong> for “.loaded-block” to appear.  \n4. Generating a <strong>screenshot</strong> &amp; <strong>PDF</strong> of the final page.  <p></p>\n<h2 id=\"9-best-practices\">9. Best Practices</h2>\n<p>1. <strong>Use <code>BrowserConfig</code> for global browser</strong> settings (headless, user agent).  \n2. <strong>Use <code>CrawlerRunConfig</code></strong> to handle the <strong>specific</strong> crawl needs: content filtering, caching, JS, screenshot, extraction, etc.  \n4. <strong>Limit</strong> large concurrency (<code>semaphore_count</code>) if the site or your system can’t handle it.  \n5. For dynamic pages, set <code>js_code</code> or <code>scan_full_page</code> so you load all content.</p>\n<h2 id=\"10-conclusion\">10. Conclusion</h2>\n<p>All parameters that used to be direct arguments to <code>arun()</code> now belong in <strong><code>CrawlerRunConfig</code></strong>. This approach:\n- Makes code <strong>clearer</strong> and <strong>more maintainable</strong>.  </p>\n<h1 id=\"arun_many-reference\"><code>arun_many(...)</code> Reference</h1>\n<blockquote>\n<p><strong>Note</strong>: This function is very similar to <a href=\"./arun.md\"><code>arun()</code></a> but focused on <strong>concurrent</strong> or <strong>batch</strong> crawling. If you’re unfamiliar with <code>arun()</code> usage, please read that doc first, then review this for differences.</p>\n</blockquote>\n<h2 id=\"function-signature\">Function Signature</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">arun_many</span>(<span class=\"hljs-params\">\n    urls: <span class=\"hljs-type\">Union</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>], <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Any</span>]],\n    config: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[CrawlerRunConfig, <span class=\"hljs-type\">List</span>[CrawlerRunConfig]]] = <span class=\"hljs-literal\">None</span>,\n    dispatcher: <span class=\"hljs-type\">Optional</span>[BaseDispatcher] = <span class=\"hljs-literal\">None</span>,\n    ...\n</span>) -&gt; <span class=\"hljs-type\">Union</span>[<span class=\"hljs-type\">List</span>[CrawlResult], AsyncGenerator[CrawlResult, <span class=\"hljs-literal\">None</span>]]:\n    <span class=\"hljs-string\">\"\"\"\n    Crawl multiple URLs concurrently or in batches.\n\n    :param urls: A list of URLs (or tasks) to crawl.\n    :param config: (Optional) Either:\n        - A single `CrawlerRunConfig` applying to all URLs\n        - A list of `CrawlerRunConfig` objects with url_matcher patterns\n    :param dispatcher: (Optional) A concurrency controller (e.g. MemoryAdaptiveDispatcher).\n    ...\n    :return: Either a list of `CrawlResult` objects, or an async generator if streaming is enabled.\n    \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"differences-from-arun\">Differences from <code>arun()</code></h2>\n<p>1. <strong>Multiple URLs</strong>:<br>\n   - Instead of crawling a single URL, you pass a list of them (strings or tasks).  \n   - The function returns either a <strong>list</strong> of <code>CrawlResult</code> or an <strong>async generator</strong> if streaming is enabled.\n2. <strong>Concurrency &amp; Dispatchers</strong>:<br>\n   - <strong><code>dispatcher</code></strong> param allows advanced concurrency control.  \n   - If omitted, a default dispatcher (like <code>MemoryAdaptiveDispatcher</code>) is used internally.  \n3. <strong>Streaming Support</strong>:<br>\n   - Enable streaming by setting <code>stream=True</code> in your <code>CrawlerRunConfig</code>.\n   - When streaming, use <code>async for</code> to process results as they become available.\n4. <strong>Parallel</strong> Execution<strong>:<br>\n   - <code>arun_many()</code> can run multiple requests concurrently under the hood.  \n   - Each <code>CrawlResult</code> might also include a </strong><code>dispatch_result</code>** with concurrency details (like memory usage, start/end times).</p>\n<h3 id=\"basic-example-batch-mode\">Basic Example (Batch Mode)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Minimal usage: The default dispatcher will be used</span>\nresults = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>],\n    config=CrawlerRunConfig(stream=<span class=\"hljs-literal\">False</span>)  <span class=\"hljs-comment\"># Default behavior</span>\n)\n\n<span class=\"hljs-keyword\">for</span> res <span class=\"hljs-keyword\">in</span> results:\n    <span class=\"hljs-keyword\">if</span> res.success:\n        <span class=\"hljs-built_in\">print</span>(res.url, <span class=\"hljs-string\">\"crawled OK!\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Failed:\"</span>, res.url, <span class=\"hljs-string\">\"-\"</span>, res.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"streaming-example\">Streaming Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    stream=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    cache_mode=CacheMode.BYPASS\n)\n\n<span class=\"hljs-comment\"># Process results as they complete</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>, <span class=\"hljs-string\">\"https://site3.com\"</span>],\n    config=config\n):\n    <span class=\"hljs-keyword\">if</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Just completed: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n        <span class=\"hljs-comment\"># Process each result immediately</span>\n        process_result(result)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"with-a-custom-dispatcher\">With a Custom Dispatcher</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">dispatcher = MemoryAdaptiveDispatcher(\n    memory_threshold_percent=70.0,\n    max_session_permit=10\n)\nresults = await crawler.arun_many(\n    urls=[<span class=\"hljs-string\">\"https://site1.com\"</span>, <span class=\"hljs-string\">\"https://site2.com\"</span>, <span class=\"hljs-string\">\"https://site3.com\"</span>],\n    config=my_run_config,\n    dispatcher=dispatcher\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"url-specific-configurations\">URL-Specific Configurations</h3>\n<p>Instead of using one config for all URLs, provide a list of configs with <code>url_matcher</code> patterns:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, MatchMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.extraction_strategy <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-comment\"># PDF files - specialized extraction</span>\npdf_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>,\n    scraping_strategy=PDFContentScrapingStrategy()\n)\n\n<span class=\"hljs-comment\"># Blog/article pages - content filtering</span>\nblog_config = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"*/blog/*\"</span>, <span class=\"hljs-string\">\"*/article/*\"</span>, <span class=\"hljs-string\">\"*python.org*\"</span>],\n    markdown_generator=DefaultMarkdownGenerator(\n        content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.48</span>)\n    )\n)\n\n<span class=\"hljs-comment\"># Dynamic pages - JavaScript execution</span>\ngithub_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'github.com'</span> <span class=\"hljs-keyword\">in</span> url,\n    js_code=<span class=\"hljs-string\">\"window.scrollTo(0, 500);\"</span>\n)\n\n<span class=\"hljs-comment\"># API endpoints - JSON extraction</span>\napi_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'api'</span> <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">or</span> url.endswith(<span class=\"hljs-string\">'.json'</span>),\n    <span class=\"hljs-comment\"># Custome settings for JSON extraction</span>\n)\n\n<span class=\"hljs-comment\"># Default fallback config</span>\ndefault_config = CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher means it never matches except as fallback</span>\n\n<span class=\"hljs-comment\"># Pass the list of configs - first match wins!</span>\nresults = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n    urls=[\n        <span class=\"hljs-string\">\"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf\"</span>,  <span class=\"hljs-comment\"># → pdf_config</span>\n        <span class=\"hljs-string\">\"https://blog.python.org/\"</span>,  <span class=\"hljs-comment\"># → blog_config</span>\n        <span class=\"hljs-string\">\"https://github.com/microsoft/playwright\"</span>,  <span class=\"hljs-comment\"># → github_config</span>\n        <span class=\"hljs-string\">\"https://httpbin.org/json\"</span>,  <span class=\"hljs-comment\"># → api_config</span>\n        <span class=\"hljs-string\">\"https://example.com/\"</span>  <span class=\"hljs-comment\"># → default_config</span>\n    ],\n    config=[pdf_config, blog_config, github_config, api_config, default_config]\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong>String patterns</strong>: <code>\"*.pdf\"</code>, <code>\"*/blog/*\"</code>, <code>\"*python.org*\"</code>\n- <strong>Function matchers</strong>: <code>lambda url: 'api' in url</code>\n- <strong>Mixed patterns</strong>: Combine strings and functions with <code>MatchMode.OR</code> or <code>MatchMode.AND</code>\n- <strong>First match wins</strong>: Configs are evaluated in order\n- <code>dispatch_result</code> in each <code>CrawlResult</code> (if using concurrency) can hold memory and timing info.  \n- <strong>Important</strong>: Always include a default config (without <code>url_matcher</code>) as the last item if you want to handle all URLs. Otherwise, unmatched URLs will fail.<p></p>\n<h3 id=\"return-value\">Return Value</h3>\n<p>Either a <strong>list</strong> of <a href=\"./crawl-result.md\"><code>CrawlResult</code></a> objects, or an <strong>async generator</strong> if streaming is enabled. You can iterate to check <code>result.success</code> or read each item’s <code>extracted_content</code>, <code>markdown</code>, or <code>dispatch_result</code>.</p>\n<h2 id=\"dispatcher-reference\">Dispatcher Reference</h2>\n<ul>\n<li><strong><code>MemoryAdaptiveDispatcher</code></strong>: Dynamically manages concurrency based on system memory usage.  </li>\n<li><strong><code>SemaphoreDispatcher</code></strong>: Fixed concurrency limit, simpler but less adaptive.  </li>\n</ul>\n<h2 id=\"common-pitfalls\">Common Pitfalls</h2>\n<p>3. <strong>Error Handling</strong>: Each <code>CrawlResult</code> might fail for different reasons—always check <code>result.success</code> or the <code>error_message</code> before proceeding.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Use <code>arun_many()</code> when you want to <strong>crawl multiple URLs</strong> simultaneously or in controlled parallel tasks. If you need advanced concurrency features (like memory-based adaptive throttling or complex rate-limiting), provide a <strong>dispatcher</strong>. Each result is a standard <code>CrawlResult</code>, possibly augmented with concurrency stats (<code>dispatch_result</code>) for deeper inspection. For more details on concurrency logic and dispatchers, see the <a href=\"../advanced/multi-url-crawling.md\">Advanced Multi-URL Crawling</a> docs.</p>\n<h1 id=\"crawlresult-reference\"><code>CrawlResult</code> Reference</h1>\n<p>The <strong><code>CrawlResult</code></strong> class encapsulates everything returned after a single crawl operation. It provides the <strong>raw or processed content</strong>, details on links and media, plus optional metadata (like screenshots, PDFs, or extracted JSON).\n<strong>Location</strong>: <code>crawl4ai/crawler/models.py</code> (for reference)\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CrawlResult</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    url: <span class=\"hljs-built_in\">str</span>\n    html: <span class=\"hljs-built_in\">str</span>\n    success: <span class=\"hljs-built_in\">bool</span>\n    cleaned_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    fit_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>  <span class=\"hljs-comment\"># Preprocessed HTML optimized for extraction</span>\n    media: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    links: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    downloaded_files: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>]] = <span class=\"hljs-literal\">None</span>\n    screenshot: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    pdf : <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">bytes</span>] = <span class=\"hljs-literal\">None</span>\n    mhtml: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    markdown: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, MarkdownGenerationResult]] = <span class=\"hljs-literal\">None</span>\n    extracted_content: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    metadata: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    error_message: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    session_id: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    response_headers: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    status_code: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">int</span>] = <span class=\"hljs-literal\">None</span>\n    ssl_certificate: <span class=\"hljs-type\">Optional</span>[SSLCertificate] = <span class=\"hljs-literal\">None</span>\n    dispatch_result: <span class=\"hljs-type\">Optional</span>[DispatchResult] = <span class=\"hljs-literal\">None</span>\n    ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"1-basic-crawl-info\">1. Basic Crawl Info</h2>\n<h3 id=\"11-url-str\">1.1 <strong><code>url</code></strong> <em>(str)</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-built_in\">print</span>(result.url)  <span class=\"hljs-comment\"># e.g., \"https://example.com/\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"12-success-bool\">1.2 <strong><code>success</code></strong> <em>(bool)</em></h3>\n<p><strong>What</strong>: <code>True</code> if the crawl pipeline ended without major errors; <code>False</code> otherwise.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"13-status_code-optionalint\">1.3 <strong><code>status_code</code></strong> <em>(Optional[int])</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.status_code == 404:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Page not found!\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"14-error_message-optionalstr\">1.4 <strong><code>error_message</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: If <code>success=False</code>, a textual description of the failure.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> not result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"15-session_id-optionalstr\">1.5 <strong><code>session_id</code></strong> <em>(Optional[str])</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># If you used session_id=\"login_session\" in CrawlerRunConfig, see it here:</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Session:\"</span>, result.session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"16-response_headers-optionaldict\">1.6 <strong><code>response_headers</code></strong> <em>(Optional[dict])</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.response_headers</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Server:\"</span>, result.response_headers.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"Server\"</span>, <span class=\"hljs-string\">\"Unknown\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"17-ssl_certificate-optionalsslcertificate\">1.7 <strong><code>ssl_certificate</code></strong> <em>(Optional[SSLCertificate])</em></h3>\n<p><strong>What</strong>: If <code>fetch_ssl_certificate=True</code> in your CrawlerRunConfig, <strong><code>result.ssl_certificate</code></strong> contains a  <a href=\"../advanced/ssl-certificate.md\"><strong><code>SSLCertificate</code></strong></a> object describing the site's certificate. You can export the cert in multiple formats (PEM/DER/JSON) or access its properties like <code>issuer</code>, \n <code>subject</code>, <code>valid_from</code>, <code>valid_until</code>, etc. \n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.ssl_certificate:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Issuer:\"</span>, result.ssl_certificate.issuer)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"2-raw-cleaned-content\">2. Raw / Cleaned Content</h2>\n<h3 id=\"21-html-str\">2.1 <strong><code>html</code></strong> <em>(str)</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Possibly large</span>\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.html))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"22-cleaned_html-optionalstr\">2.2 <strong><code>cleaned_html</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: A sanitized HTML version—scripts, styles, or excluded tags are removed based on your <code>CrawlerRunConfig</code>.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-built_in\">print</span>(result.cleaned_html[:500])  <span class=\"hljs-comment\"># Show a snippet</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"3-markdown-fields\">3. Markdown Fields</h2>\n<h3 id=\"31-the-markdown-generation-approach\">3.1 The Markdown Generation Approach</h3>\n<ul>\n<li><strong>Raw</strong> markdown  </li>\n<li><strong>Links as citations</strong> (with a references section)  </li>\n<li><strong>Fit</strong> markdown if a <strong>content filter</strong> is used (like Pruning or BM25)\n<strong><code>MarkdownGenerationResult</code></strong> includes:</li>\n<li><strong><code>raw_markdown</code></strong> <em>(str)</em>: The full HTML→Markdown conversion.  </li>\n<li><strong><code>markdown_with_citations</code></strong> <em>(str)</em>: Same markdown, but with link references as academic-style citations.  </li>\n<li><strong><code>references_markdown</code></strong> <em>(str)</em>: The reference list or footnotes at the end.  </li>\n<li><strong><code>fit_markdown</code></strong> <em>(Optional[str])</em>: If content filtering (Pruning/BM25) was applied, the filtered \"fit\" text.  </li>\n<li><strong><code>fit_html</code></strong> <em>(Optional[str])</em>: The HTML that led to <code>fit_markdown</code>.\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.markdown:\n    md_res = result.markdown\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw MD:\"</span>, md_res.raw_markdown[:300])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Citations MD:\"</span>, md_res.markdown_with_citations[:300])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"References:\"</span>, md_res.references_markdown)\n    <span class=\"hljs-keyword\">if</span> md_res.fit_markdown:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Pruned text:\"</span>, md_res.fit_markdown[:300])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h3 id=\"32-markdown-optionalunionstr-markdowngenerationresult\">3.2 <strong><code>markdown</code></strong> <em>(Optional[Union[str, MarkdownGenerationResult]])</em></h3>\n<p><strong>What</strong>: Holds the <code>MarkdownGenerationResult</code>.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-scss\"><span class=\"hljs-built_in\">print</span>(result.markdown.raw_markdown[:<span class=\"hljs-number\">200</span>])\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_html)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Important</strong>: \"Fit\" content (in <code>fit_markdown</code>/<code>fit_html</code>) exists in result.markdown, only if you used a <strong>filter</strong> (like <strong>PruningContentFilter</strong> or <strong>BM25ContentFilter</strong>) within a <code>MarkdownGenerationStrategy</code>.<p></p>\n<h2 id=\"4-media-links\">4. Media &amp; Links</h2>\n<h3 id=\"41-media-dictstr-listdict\">4.1 <strong><code>media</code></strong> <em>(Dict[str, List[Dict]])</em></h3>\n<p><strong>What</strong>: Contains info about discovered images, videos, or audio. Typically keys: <code>\"images\"</code>, <code>\"videos\"</code>, <code>\"audios\"</code>.<br>\n- <code>src</code> <em>(str)</em>: Media URL<br>\n- <code>alt</code> or <code>title</code> <em>(str)</em>: Descriptive text<br>\n- <code>score</code> <em>(float)</em>: Relevance score if the crawler's heuristic found it \"important\"<br>\n- <code>desc</code> or <code>description</code> <em>(Optional[str])</em>: Additional context extracted from surrounding text<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">images = result.media.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"images\"</span>, [])\n<span class=\"hljs-keyword\">for</span> img <span class=\"hljs-keyword\">in</span> images:\n    <span class=\"hljs-keyword\">if</span> img.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"score\"</span>, <span class=\"hljs-number\">0</span>) &gt; <span class=\"hljs-number\">5</span>:\n        print(<span class=\"hljs-string\">\"High-value image:\"</span>, img[<span class=\"hljs-string\">\"src\"</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"42-links-dictstr-listdict\">4.2 <strong><code>links</code></strong> <em>(Dict[str, List[Dict]])</em></h3>\n<p><strong>What</strong>: Holds internal and external link data. Usually two keys: <code>\"internal\"</code> and <code>\"external\"</code>.<br>\n- <code>href</code> <em>(str)</em>: The link target<br>\n- <code>text</code> <em>(str)</em>: Link text<br>\n- <code>title</code> <em>(str)</em>: Title attribute<br>\n- <code>context</code> <em>(str)</em>: Surrounding text snippet<br>\n- <code>domain</code> <em>(str)</em>: If external, the domain\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links[<span class=\"hljs-string\">\"internal\"</span>]:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal link to <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span> with text <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'text'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"5-additional-fields\">5. Additional Fields</h2>\n<h3 id=\"51-extracted_content-optionalstr\">5.1 <strong><code>extracted_content</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: If you used <strong><code>extraction_strategy</code></strong> (CSS, LLM, etc.), the structured output (JSON).<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.extracted_content</span>:\n    data = json.<span class=\"hljs-built_in\">loads</span>(result.extracted_content)\n    <span class=\"hljs-built_in\">print</span>(data)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"52-downloaded_files-optionalliststr\">5.2 <strong><code>downloaded_files</code></strong> <em>(Optional[List[str]])</em></h3>\n<p><strong>What</strong>: If <code>accept_downloads=True</code> in your <code>BrowserConfig</code> + <code>downloads_path</code>, lists local file paths for downloaded items.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-keyword\">if</span> result.downloaded_files:\n    <span class=\"hljs-keyword\">for</span> file_path <span class=\"hljs-keyword\">in</span> result.downloaded_files:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Downloaded:\"</span>, file_path)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"53-screenshot-optionalstr\">5.3 <strong><code>screenshot</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: Base64-encoded screenshot if <code>screenshot=True</code> in <code>CrawlerRunConfig</code>.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> base64\n<span class=\"hljs-keyword\">if</span> result.screenshot:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.png\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(base64.b64decode(result.screenshot))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"54-pdf-optionalbytes\">5.4 <strong><code>pdf</code></strong> <em>(Optional[bytes])</em></h3>\n<p><strong>What</strong>: Raw PDF bytes if <code>pdf=True</code> in <code>CrawlerRunConfig</code>.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.pdf:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.pdf\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(result.pdf)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"55-mhtml-optionalstr\">5.5 <strong><code>mhtml</code></strong> <em>(Optional[str])</em></h3>\n<p><strong>What</strong>: MHTML snapshot of the page if <code>capture_mhtml=True</code> in <code>CrawlerRunConfig</code>. MHTML (MIME HTML) format preserves the entire web page with all its resources (CSS, images, scripts, etc.) in a single file.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.mhtml:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.mhtml\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(result.mhtml)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"56-metadata-optionaldict\">5.6 <strong><code>metadata</code></strong> <em>(Optional[dict])</em></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.metadata</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Title:\"</span>, result.metadata.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"title\"</span>))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Author:\"</span>, result.metadata.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"author\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"6-dispatch_result-optional\">6. <code>dispatch_result</code> (optional)</h2>\n<p>A <code>DispatchResult</code> object providing additional concurrency and resource usage information when crawling URLs in parallel (e.g., via <code>arun_many()</code> with custom dispatchers). It contains:\n- <strong><code>task_id</code></strong>: A unique identifier for the parallel task.\n- <strong><code>memory_usage</code></strong> (float): The memory (in MB) used at the time of completion.\n- <strong><code>peak_memory</code></strong> (float): The peak memory usage (in MB) recorded during the task's execution.\n- <strong><code>start_time</code></strong> / <strong><code>end_time</code></strong> (datetime): Time range for this crawling task.\n- <strong><code>error_message</code></strong> (str): Any dispatcher- or concurrency-related error encountered.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Example usage:</span>\n<span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n    <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.dispatch_result:\n        dr = result.dispatch_result\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"URL: <span class=\"hljs-subst\">{result.url}</span>, Task ID: <span class=\"hljs-subst\">{dr.task_id}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Memory: <span class=\"hljs-subst\">{dr.memory_usage:<span class=\"hljs-number\">.1</span>f}</span> MB (Peak: <span class=\"hljs-subst\">{dr.peak_memory:<span class=\"hljs-number\">.1</span>f}</span> MB)\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Duration: <span class=\"hljs-subst\">{dr.end_time - dr.start_time}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<blockquote>\n<p><strong>Note</strong>: This field is typically populated when using <code>arun_many(...)</code> alongside a <strong>dispatcher</strong> (e.g., <code>MemoryAdaptiveDispatcher</code> or <code>SemaphoreDispatcher</code>). If no concurrency or dispatcher is used, <code>dispatch_result</code> may remain <code>None</code>. </p>\n</blockquote>\n<h2 id=\"7-network-requests-console-messages\">7. Network Requests &amp; Console Messages</h2>\n<p>When you enable network and console message capturing in <code>CrawlerRunConfig</code> using <code>capture_network_requests=True</code> and <code>capture_console_messages=True</code>, the <code>CrawlResult</code> will include these fields:</p>\n<h3 id=\"71-network_requests-optionallistdictstr-any\">7.1 <strong><code>network_requests</code></strong> <em>(Optional[List[Dict[str, Any]]])</em></h3>\n<ul>\n<li>Each item has an <code>event_type</code> field that can be <code>\"request\"</code>, <code>\"response\"</code>, or <code>\"request_failed\"</code>.</li>\n<li>Request events include <code>url</code>, <code>method</code>, <code>headers</code>, <code>post_data</code>, <code>resource_type</code>, and <code>is_navigation_request</code>.</li>\n<li>Response events include <code>url</code>, <code>status</code>, <code>status_text</code>, <code>headers</code>, and <code>request_timing</code>.</li>\n<li>Failed request events include <code>url</code>, <code>method</code>, <code>resource_type</code>, and <code>failure_text</code>.</li>\n<li>All events include a <code>timestamp</code> field.\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.network_requests:\n    <span class=\"hljs-comment\"># Count different types of events</span>\n    requests = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request\"</span>]\n    responses = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"response\"</span>]\n    failures = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> result.network_requests <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"event_type\"</span>) == <span class=\"hljs-string\">\"request_failed\"</span>]\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Captured <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(requests)}</span> requests, <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(responses)}</span> responses, and <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(failures)}</span> failures\"</span>)\n\n    <span class=\"hljs-comment\"># Analyze API calls</span>\n    api_calls = [r <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> requests <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"api\"</span> <span class=\"hljs-keyword\">in</span> r.get(<span class=\"hljs-string\">\"url\"</span>, <span class=\"hljs-string\">\"\"</span>)]\n\n    <span class=\"hljs-comment\"># Identify failed resources</span>\n    <span class=\"hljs-keyword\">for</span> failure <span class=\"hljs-keyword\">in</span> failures:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to load: <span class=\"hljs-subst\">{failure.get(<span class=\"hljs-string\">'url'</span>)}</span> - <span class=\"hljs-subst\">{failure.get(<span class=\"hljs-string\">'failure_text'</span>)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h3 id=\"72-console_messages-optionallistdictstr-any\">7.2 <strong><code>console_messages</code></strong> <em>(Optional[List[Dict[str, Any]]])</em></h3>\n<ul>\n<li>Each item has a <code>type</code> field indicating the message type (e.g., <code>\"log\"</code>, <code>\"error\"</code>, <code>\"warning\"</code>, etc.).</li>\n<li>The <code>text</code> field contains the actual message text.</li>\n<li>Some messages include <code>location</code> information (URL, line, column).</li>\n<li>All messages include a <code>timestamp</code> field.\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.console_messages:\n    <span class=\"hljs-comment\"># Count messages by type</span>\n    message_types = {}\n    <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n        msg_type = msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)\n        message_types[msg_type] = message_types.get(msg_type, <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Message type counts: <span class=\"hljs-subst\">{message_types}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Display errors (which are usually most important)</span>\n    <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n        <span class=\"hljs-keyword\">if</span> msg.get(<span class=\"hljs-string\">\"type\"</span>) == <span class=\"hljs-string\">\"error\"</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{msg.get(<span class=\"hljs-string\">'text'</span>)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h2 id=\"8-example-accessing-everything\">8. Example: Accessing Everything</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">handle_result</span>(<span class=\"hljs-params\">result: CrawlResult</span>):\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl error:\"</span>, result.error_message)\n        <span class=\"hljs-keyword\">return</span>\n\n    <span class=\"hljs-comment\"># Basic info</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawled URL:\"</span>, result.url)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Status code:\"</span>, result.status_code)\n\n    <span class=\"hljs-comment\"># HTML</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Original HTML size:\"</span>, <span class=\"hljs-built_in\">len</span>(result.html))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML size:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html <span class=\"hljs-keyword\">or</span> <span class=\"hljs-string\">\"\"</span>))\n\n    <span class=\"hljs-comment\"># Markdown output</span>\n    <span class=\"hljs-keyword\">if</span> result.markdown:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown:\"</span>, result.markdown.raw_markdown[:<span class=\"hljs-number\">300</span>])\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Citations Markdown:\"</span>, result.markdown.markdown_with_citations[:<span class=\"hljs-number\">300</span>])\n        <span class=\"hljs-keyword\">if</span> result.markdown.fit_markdown:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown:\"</span>, result.markdown.fit_markdown[:<span class=\"hljs-number\">200</span>])\n\n    <span class=\"hljs-comment\"># Media &amp; Links</span>\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"images\"</span> <span class=\"hljs-keyword\">in</span> result.media:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Image count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.media[<span class=\"hljs-string\">\"images\"</span>]))\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"internal\"</span> <span class=\"hljs-keyword\">in</span> result.links:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal link count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links[<span class=\"hljs-string\">\"internal\"</span>]))\n\n    <span class=\"hljs-comment\"># Extraction strategy result</span>\n    <span class=\"hljs-keyword\">if</span> result.extracted_content:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Structured data:\"</span>, result.extracted_content)\n\n    <span class=\"hljs-comment\"># Screenshot/PDF/MHTML</span>\n    <span class=\"hljs-keyword\">if</span> result.screenshot:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n    <span class=\"hljs-keyword\">if</span> result.pdf:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"PDF bytes length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.pdf))\n    <span class=\"hljs-keyword\">if</span> result.mhtml:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"MHTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.mhtml))\n\n    <span class=\"hljs-comment\"># Network and console capturing</span>\n    <span class=\"hljs-keyword\">if</span> result.network_requests:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Network requests captured: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.network_requests)}</span>\"</span>)\n        <span class=\"hljs-comment\"># Analyze request types</span>\n        req_types = {}\n        <span class=\"hljs-keyword\">for</span> req <span class=\"hljs-keyword\">in</span> result.network_requests:\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"resource_type\"</span> <span class=\"hljs-keyword\">in</span> req:\n                req_types[req[<span class=\"hljs-string\">\"resource_type\"</span>]] = req_types.get(req[<span class=\"hljs-string\">\"resource_type\"</span>], <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Resource types: <span class=\"hljs-subst\">{req_types}</span>\"</span>)\n\n    <span class=\"hljs-keyword\">if</span> result.console_messages:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Console messages captured: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.console_messages)}</span>\"</span>)\n        <span class=\"hljs-comment\"># Count by message type</span>\n        msg_types = {}\n        <span class=\"hljs-keyword\">for</span> msg <span class=\"hljs-keyword\">in</span> result.console_messages:\n            msg_types[msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)] = msg_types.get(msg.get(<span class=\"hljs-string\">\"type\"</span>, <span class=\"hljs-string\">\"unknown\"</span>), <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Message types: <span class=\"hljs-subst\">{msg_types}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"9-key-points-future\">9. Key Points &amp; Future</h2>\n<p>1. <strong>Deprecated legacy properties of CrawlResult</strong><br>\n   - <code>markdown_v2</code> - Removed in v0.5. Accessing it raises <code>AttributeError</code>. Use <code>markdown</code>.\n   - <code>fit_markdown</code> and <code>fit_html</code> - Removed as top-level <code>CrawlResult</code> properties in v0.5. Use <code>result.markdown.fit_markdown</code> and <code>result.markdown.fit_html</code>.\n2. <strong>Fit Content</strong><br>\n   - <strong><code>fit_markdown</code></strong> and <strong><code>fit_html</code></strong> appear in MarkdownGenerationResult, only if you used a content filter (like <strong>PruningContentFilter</strong> or <strong>BM25ContentFilter</strong>) inside your <strong>MarkdownGenerationStrategy</strong> or set them directly.<br>\n   - If no filter is used, they remain <code>None</code>.\n3. <strong>References &amp; Citations</strong><br>\n   - If you enable link citations in your <code>DefaultMarkdownGenerator</code> (<code>options={\"citations\": True}</code>), you’ll see <code>markdown_with_citations</code> plus a <strong><code>references_markdown</code></strong> block. This helps large language models or academic-like referencing.\n4. <strong>Links &amp; Media</strong><br>\n   - <code>links[\"internal\"]</code> and <code>links[\"external\"]</code> group discovered anchors by domain.<br>\n   - <code>media[\"images\"]</code> / <code>[\"videos\"]</code> / <code>[\"audios\"]</code> store extracted media elements with optional scoring or context.\n5. <strong>Error Cases</strong><br>\n   - If <code>success=False</code>, check <code>error_message</code> (e.g., timeouts, invalid URLs).<br>\n   - <code>status_code</code> might be <code>None</code> if we failed before an HTTP response.\nUse <strong><code>CrawlResult</code></strong> to glean all final outputs and feed them into your data pipelines, AI models, or archives. With the synergy of a properly configured <strong>BrowserConfig</strong> and <strong>CrawlerRunConfig</strong>, the crawler can produce robust, structured results here in <strong><code>CrawlResult</code></strong>.</p>\n<h1 id=\"configuration\">Configuration</h1>\n<h1 id=\"browser-crawler-llm-configuration-quick-overview\">Browser, Crawler &amp; LLM Configuration (Quick Overview)</h1>\n<p>Crawl4AI's flexibility stems from two key classes:\n1. <strong><code>BrowserConfig</code></strong> – Dictates <strong>how</strong> the browser is launched and behaves (e.g., headless or visible, proxy, user agent).<br>\n2. <strong><code>CrawlerRunConfig</code></strong> – Dictates <strong>how</strong> each <strong>crawl</strong> operates (e.g., caching, extraction, timeouts, JavaScript code to run, etc.).<br>\n3. <strong><code>LLMConfig</code></strong> - Dictates <strong>how</strong> LLM providers are configured. (model, api token, base url, temperature etc.)\nIn most examples, you create <strong>one</strong> <code>BrowserConfig</code> for the entire crawler session, then pass a <strong>fresh</strong> or re-used <code>CrawlerRunConfig</code> whenever you call <code>arun()</code>. This tutorial shows the most commonly used parameters. If you need advanced or rarely used fields, see the <a href=\"../api/parameters.md\">Configuration Parameters</a>.</p>\n<h2 id=\"1-browserconfig-essentials\">1. BrowserConfig Essentials</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">BrowserConfig</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n        headless=<span class=\"hljs-literal\">True</span>,\n        proxy_config=<span class=\"hljs-literal\">None</span>,\n        viewport_width=<span class=\"hljs-number\">1080</span>,\n        viewport_height=<span class=\"hljs-number\">600</span>,\n        verbose=<span class=\"hljs-literal\">True</span>,\n        use_persistent_context=<span class=\"hljs-literal\">False</span>,\n        user_data_dir=<span class=\"hljs-literal\">None</span>,\n        cookies=<span class=\"hljs-literal\">None</span>,\n        headers=<span class=\"hljs-literal\">None</span>,\n        user_agent=<span class=\"hljs-literal\">None</span>,\n        text_mode=<span class=\"hljs-literal\">False</span>,\n        light_mode=<span class=\"hljs-literal\">False</span>,\n        avoid_ads=<span class=\"hljs-literal\">False</span>,\n        avoid_css=<span class=\"hljs-literal\">False</span>,\n        extra_args=<span class=\"hljs-literal\">None</span>,\n        enable_stealth=<span class=\"hljs-literal\">False</span>,\n        <span class=\"hljs-comment\"># ... other advanced parameters omitted here</span>\n    </span>):\n        ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"key-fields-to-note\">Key Fields to Note</h3>\n<ol>\n<li><strong><code>browser_type</code></strong>  </li>\n<li>Options: <code>\"chromium\"</code>, <code>\"firefox\"</code>, or <code>\"webkit\"</code>.  </li>\n<li>Defaults to <code>\"chromium\"</code>.  </li>\n<li>If you need a different engine, specify it here.</li>\n<li><strong><code>headless</code></strong>  </li>\n<li><code>True</code>: Runs the browser in headless mode (invisible browser).  </li>\n<li><code>False</code>: Runs the browser in visible mode, which helps with debugging.</li>\n<li><strong><code>proxy_config</code></strong>  </li>\n<li>A dictionary with fields like:<br>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-attr\">\"server\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span><span class=\"hljs-punctuation\">,</span> \n    <span class=\"hljs-attr\">\"username\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span><span class=\"hljs-punctuation\">,</span> \n    <span class=\"hljs-attr\">\"password\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li>Leave as <code>None</code> if a proxy is not required.</li>\n<li><strong><code>viewport_width</code> &amp; <code>viewport_height</code></strong>:  </li>\n<li>The initial window size.  </li>\n<li>Some sites behave differently with smaller or bigger viewports.</li>\n<li><strong><code>verbose</code></strong>:  </li>\n<li>If <code>True</code>, prints extra logs.  </li>\n<li>Handy for debugging.</li>\n<li><strong><code>use_persistent_context</code></strong>:  </li>\n<li>If <code>True</code>, uses a <strong>persistent</strong> browser profile, storing cookies/local storage across runs.  </li>\n<li>Typically also set <code>user_data_dir</code> to point to a folder.</li>\n<li><strong><code>cookies</code></strong> &amp; <strong><code>headers</code></strong>:  </li>\n<li>E.g. <code>cookies=[{\"name\": \"session\", \"value\": \"abc123\", \"domain\": \"example.com\"}]</code>.</li>\n<li><strong><code>user_agent</code></strong>:  </li>\n<li>Custom User-Agent string. If <code>None</code>, a default is used.  </li>\n<li>You can also set <code>user_agent_mode=\"random\"</code> for randomization (if you want to fight bot detection).</li>\n<li><strong><code>text_mode</code></strong> &amp; <strong><code>light_mode</code></strong>:</li>\n<li><code>text_mode=True</code> disables images, possibly speeding up text-only crawls.</li>\n<li><code>light_mode=True</code> turns off certain background features for performance.</li>\n<li><strong><code>avoid_ads</code></strong> &amp; <strong><code>avoid_css</code></strong>:<ul>\n<li><code>avoid_ads=True</code> blocks requests to common ad and tracker domains (Google Analytics, DoubleClick, Facebook, Hotjar, etc.) at the browser context level. Reduces network overhead and memory usage.</li>\n<li><code>avoid_css=True</code> blocks loading of CSS files (<code>.css</code>, <code>.less</code>, <code>.scss</code>, <code>.sass</code>), useful when you only need text content and want faster, leaner crawls.</li>\n<li>Both default to <code>False</code> (opt-in). Can be combined with each other and with <code>text_mode</code>.</li>\n</ul>\n</li>\n<li><strong><code>extra_args</code></strong>:<ul>\n<li>Additional flags for the underlying browser.</li>\n<li>E.g. <code>[\"--disable-extensions\"]</code>.</li>\n</ul>\n</li>\n<li><strong><code>enable_stealth</code></strong>:<ul>\n<li>If <code>True</code>, enables stealth mode using playwright-stealth.</li>\n<li>Modifies browser fingerprints to avoid basic bot detection.</li>\n<li>Default is <code>False</code>. Recommended for sites with bot protection.</li>\n</ul>\n</li>\n</ol>\n<h3 id=\"helper-methods\">Helper Methods</h3>\n<p>Both configuration classes provide a <code>clone()</code> method to create modified copies:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Create a base browser config</span>\nbase_browser <span class=\"hljs-punctuation\">=</span> BrowserConfig<span class=\"hljs-punctuation\">(</span>\n    browser_type<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"chromium\"</span>,\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    text_mode<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Create a visible browser config for debugging</span>\ndebug_browser <span class=\"hljs-punctuation\">=</span> base_browser.clone<span class=\"hljs-punctuation\">(</span>\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>,\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nbrowser_conf = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"firefox\"</span>,\n    headless=<span class=\"hljs-literal\">False</span>,\n    text_mode=<span class=\"hljs-literal\">True</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n    <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"2-crawlerrunconfig-essentials\">2. CrawlerRunConfig Essentials</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CrawlerRunConfig</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        word_count_threshold=<span class=\"hljs-number\">200</span>,\n        extraction_strategy=<span class=\"hljs-literal\">None</span>,\n        markdown_generator=<span class=\"hljs-literal\">None</span>,\n        cache_mode=<span class=\"hljs-literal\">None</span>,\n        js_code=<span class=\"hljs-literal\">None</span>,\n        wait_for=<span class=\"hljs-literal\">None</span>,\n        screenshot=<span class=\"hljs-literal\">False</span>,\n        pdf=<span class=\"hljs-literal\">False</span>,\n        capture_mhtml=<span class=\"hljs-literal\">False</span>,\n        <span class=\"hljs-comment\"># Location and Identity Parameters</span>\n        locale=<span class=\"hljs-literal\">None</span>,            <span class=\"hljs-comment\"># e.g. \"en-US\", \"fr-FR\"</span>\n        timezone_id=<span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># e.g. \"America/New_York\"</span>\n        geolocation=<span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># GeolocationConfig object</span>\n        <span class=\"hljs-comment\"># Resource Management</span>\n        enable_rate_limiting=<span class=\"hljs-literal\">False</span>,\n        rate_limit_config=<span class=\"hljs-literal\">None</span>,\n        memory_threshold_percent=<span class=\"hljs-number\">70.0</span>,\n        check_interval=<span class=\"hljs-number\">1.0</span>,\n        max_session_permit=<span class=\"hljs-number\">20</span>,\n        display_mode=<span class=\"hljs-literal\">None</span>,\n        verbose=<span class=\"hljs-literal\">True</span>,\n        stream=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Enable streaming for arun_many()</span>\n        <span class=\"hljs-comment\"># ... other advanced parameters omitted</span>\n    </span>):\n        ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"key-fields-to-note_1\">Key Fields to Note</h3>\n<ol>\n<li><strong><code>word_count_threshold</code></strong>:  </li>\n<li>The minimum word count before a block is considered.  </li>\n<li>If your site has lots of short paragraphs or items, you can lower it.</li>\n<li><strong><code>extraction_strategy</code></strong>:  </li>\n<li>Where you plug in JSON-based extraction (CSS, LLM, etc.).  </li>\n<li>If <code>None</code>, no structured extraction is done (only raw/cleaned HTML + markdown).</li>\n<li><strong><code>markdown_generator</code></strong>:  </li>\n<li>E.g., <code>DefaultMarkdownGenerator(...)</code>, controlling how HTML→Markdown conversion is done.  </li>\n<li>If <code>None</code>, a default approach is used.</li>\n<li><strong><code>cache_mode</code></strong>:  </li>\n<li>Controls caching behavior (<code>ENABLED</code>, <code>BYPASS</code>, <code>DISABLED</code>, etc.).  </li>\n<li>If <code>None</code>, defaults to some level of caching or you can specify <code>CacheMode.ENABLED</code>.</li>\n<li><strong><code>js_code</code></strong>:  </li>\n<li>A string or list of JS strings to execute.  </li>\n<li>Great for \"Load More\" buttons or user interactions.  </li>\n<li><strong><code>wait_for</code></strong>:  </li>\n<li>A CSS or JS expression to wait for before extracting content.  </li>\n<li>Common usage: <code>wait_for=\"css:.main-loaded\"</code> or <code>wait_for=\"js:() =&gt; window.loaded === true\"</code>.</li>\n<li><strong><code>screenshot</code></strong>, <strong><code>pdf</code></strong>, &amp; <strong><code>capture_mhtml</code></strong>:  </li>\n<li>If <code>True</code>, captures a screenshot, PDF, or MHTML snapshot after the page is fully loaded.  </li>\n<li>The results go to <code>result.screenshot</code> (base64), <code>result.pdf</code> (bytes), or <code>result.mhtml</code> (string).</li>\n<li><strong>Location Parameters</strong>:  </li>\n<li><strong><code>locale</code></strong>: Browser's locale (e.g., <code>\"en-US\"</code>, <code>\"fr-FR\"</code>) for language preferences</li>\n<li><strong><code>timezone_id</code></strong>: Browser's timezone (e.g., <code>\"America/New_York\"</code>, <code>\"Europe/Paris\"</code>)</li>\n<li><strong><code>geolocation</code></strong>: GPS coordinates via <code>GeolocationConfig(latitude=48.8566, longitude=2.3522)</code></li>\n<li><strong><code>verbose</code></strong>:  </li>\n<li>Logs additional runtime details.  </li>\n<li>Overlaps with the browser's verbosity if also set to <code>True</code> in <code>BrowserConfig</code>.</li>\n<li><strong><code>enable_rate_limiting</code></strong>:  </li>\n<li>If <code>True</code>, enables rate limiting for batch processing.  </li>\n<li>Requires <code>rate_limit_config</code> to be set.</li>\n<li><strong><code>memory_threshold_percent</code></strong>:  <ul>\n<li>The memory threshold (as a percentage) to monitor.  </li>\n<li>If exceeded, the crawler will pause or slow down.</li>\n</ul>\n</li>\n<li><strong><code>check_interval</code></strong>:  <ul>\n<li>The interval (in seconds) to check system resources.  </li>\n<li>Affects how often memory and CPU usage are monitored.</li>\n</ul>\n</li>\n<li><strong><code>max_session_permit</code></strong>:  <ul>\n<li>The maximum number of concurrent crawl sessions.  </li>\n<li>Helps prevent overwhelming the system.</li>\n</ul>\n</li>\n<li><strong><code>url_matcher</code></strong> &amp; <strong><code>match_mode</code></strong>:  <ul>\n<li>Enable URL-specific configurations when used with <code>arun_many()</code>.</li>\n<li>Set <code>url_matcher</code> to patterns (glob, function, or list) to match specific URLs.</li>\n<li>Use <code>match_mode</code> (OR/AND) to control how multiple patterns combine.</li>\n</ul>\n</li>\n<li><strong><code>display_mode</code></strong>:  <ul>\n<li>The display mode for progress information (<code>DETAILED</code>, <code>BRIEF</code>, etc.).  </li>\n<li>Affects how much information is printed during the crawl.</li>\n</ul>\n</li>\n</ol>\n<h3 id=\"helper-methods_1\">Helper Methods</h3>\n<p>The <code>clone()</code> method is particularly useful for creating variations of your crawler configuration:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Create a base configuration</span>\nbase_config = CrawlerRunConfig(\n    cache_mode=CacheMode.ENABLED,\n    word_count_threshold=200,\n    wait_until=<span class=\"hljs-string\">\"networkidle\"</span>\n)\n\n<span class=\"hljs-comment\"># Create variations for different use cases</span>\nstream_config = base_config.clone(\n    stream=True,  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    cache_mode=CacheMode.BYPASS\n)\n\ndebug_config = base_config.clone(\n    page_timeout=120000,  <span class=\"hljs-comment\"># Longer timeout for debugging</span>\n    verbose=True\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThe <code>clone()</code> method:\n- Creates a new instance with all the same settings\n- Updates only the specified parameters\n- Leaves the original configuration unchanged\n- Perfect for creating variations without repeating all parameters<p></p>\n<h2 id=\"3-llmconfig-essentials\">3. LLMConfig Essentials</h2>\n<h3 id=\"key-fields-to-note_2\">Key fields to note</h3>\n<ol>\n<li><strong><code>provider</code></strong>:  </li>\n<li>Which LLM provider to use. </li>\n<li>Possible values are <code>\"ollama/llama3\",\"groq/llama3-70b-8192\",\"groq/llama3-8b-8192\", \"openai/gpt-4o-mini\" ,\"openai/gpt-4o\",\"openai/o1-mini\",\"openai/o1-preview\",\"openai/o3-mini\",\"openai/o3-mini-high\",\"anthropic/claude-3-haiku-20240307\",\"anthropic/claude-3-opus-20240229\",\"anthropic/claude-3-sonnet-20240229\",\"anthropic/claude-3-5-sonnet-20240620\",\"gemini/gemini-pro\",\"gemini/gemini-1.5-pro\",\"gemini/gemini-2.0-flash\",\"gemini/gemini-2.0-flash-exp\",\"gemini/gemini-2.0-flash-lite-preview-02-05\",\"deepseek/deepseek-chat\"</code><br><em>(default: <code>\"openai/gpt-4o-mini\"</code>)</em></li>\n<li><strong><code>api_token</code></strong>:  <ul>\n<li>Optional. When not provided explicitly, api_token will be read from environment variables based on provider. For example: If a gemini model is passed as provider then,<code>\"GEMINI_API_KEY\"</code> will be read from environment variables  </li>\n<li>API token of LLM provider <br> eg: <code>api_token = \"gsk_1ClHGGJ7Lpn4WGybR7vNWGdyb3FY7zXEw3SCiy0BAVM9lL8CQv\"</code></li>\n<li>Environment variable - use with prefix \"env:\" <br> eg:<code>api_token = \"env: GROQ_API_KEY\"</code>            </li>\n</ul>\n</li>\n<li><strong><code>base_url</code></strong>:  </li>\n<li>\n<p>If your provider has a custom endpoint</p>\n</li>\n<li>\n<p><strong>Backoff controls</strong> <em>(optional)</em>:  </p>\n</li>\n<li><code>backoff_base_delay</code> <em>(default <code>2</code> seconds)</em> – how long to pause before the first retry if the provider rate-limits you.  </li>\n<li><code>backoff_max_attempts</code> <em>(default <code>3</code>)</em> – total tries for the same prompt (initial call + retries).  </li>\n<li><code>backoff_exponential_factor</code> <em>(default <code>2</code>)</em> – how quickly the pause grows between retries. A factor of 2 yields waits like 2s → 4s → 8s.  </li>\n<li>Because these plug into Crawl4AI’s retry helper, every LLM strategy automatically follows the pacing you define here.\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">llm_config = LLMConfig(\n    provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>,\n    api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>),\n    backoff_base_delay=1, <span class=\"hljs-comment\"># optional</span>\n    backoff_max_attempts=5, <span class=\"hljs-comment\"># optional</span>\n    backoff_exponential_factor=3, <span class=\"hljs-comment\"># optional</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ol>\n<h2 id=\"4-putting-it-all-together\">4. Putting It All Together</h2>\n<p>In a typical scenario, you define <strong>one</strong> <code>BrowserConfig</code> for your crawler session, then create <strong>one or more</strong> <code>CrawlerRunConfig</code> &amp; <code>LLMConfig</code> depending on each call's needs:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig, LLMContentFilter, DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1) Browser config: headless, bigger viewport, no proxy</span>\n    browser_conf = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,\n        viewport_width=<span class=\"hljs-number\">1280</span>,\n        viewport_height=<span class=\"hljs-number\">720</span>\n    )\n\n    <span class=\"hljs-comment\"># 2) Example extraction strategy</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.article\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n    extraction = JsonCssExtractionStrategy(schema)\n\n    <span class=\"hljs-comment\"># 3) Example LLM content filtering</span>\n\n    gemini_config = LLMConfig(\n        provider=<span class=\"hljs-string\">\"gemini/gemini-1.5-pro\"</span>, \n        api_token = <span class=\"hljs-string\">\"env:GEMINI_API_TOKEN\"</span>\n    )\n\n    <span class=\"hljs-comment\"># Initialize LLM filter with specific instruction</span>\n    <span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n        llm_config=gemini_config,  <span class=\"hljs-comment\"># or your preferred provider</span>\n        instruction=<span class=\"hljs-string\">\"\"\"\n        Focus on extracting the core educational content.\n        Include:\n        - Key concepts and explanations\n        - Important code examples\n        - Essential technical details\n        Exclude:\n        - Navigation elements\n        - Sidebars\n        - Footer content\n        Format the output as clean markdown with proper code blocks and headers.\n        \"\"\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">500</span>,  <span class=\"hljs-comment\"># Adjust based on your needs</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    md_generator = DefaultMarkdownGenerator(\n        content_filter=<span class=\"hljs-built_in\">filter</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># 4) Crawler run config: skip cache, use extraction</span>\n    run_conf = CrawlerRunConfig(\n        markdown_generator=md_generator,\n        extraction_strategy=extraction,\n        cache_mode=CacheMode.BYPASS,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4) Execute the crawl</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com/news\"</span>, config=run_conf)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted content:\"</span>, result.extracted_content)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"5-next-steps\">5. Next Steps</h2>\n<ul>\n<li><a href=\"../api/parameters.md\">BrowserConfig, CrawlerRunConfig &amp; LLMConfig Reference</a>  </li>\n<li><strong>Custom Hooks &amp; Auth</strong> (Inject JavaScript or handle login forms).  </li>\n<li><strong>Session Management</strong> (Re-use pages, preserve state across multiple calls).  </li>\n<li><strong>Advanced Caching</strong> (Fine-tune read/write cache modes).  </li>\n</ul>\n<h2 id=\"6-conclusion\">6. Conclusion</h2>\n<h1 id=\"1-browserconfig-controlling-the-browser\">1. <strong>BrowserConfig</strong> – Controlling the Browser</h1>\n<p><code>BrowserConfig</code> focuses on <strong>how</strong> the browser is launched and behaves. This includes headless mode, proxies, user agents, and other environment tweaks.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nbrowser_cfg = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n    headless=<span class=\"hljs-literal\">True</span>,\n    viewport_width=<span class=\"hljs-number\">1280</span>,\n    viewport_height=<span class=\"hljs-number\">720</span>,\n    proxy_config=<span class=\"hljs-string\">\"http://user:pass@proxy:8080\"</span>,\n    user_agent=<span class=\"hljs-string\">\"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 Chrome/116.0.0.0 Safari/537.36\"</span>,\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"11-parameter-highlights\">1.1 Parameter Highlights</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th><strong>Parameter</strong></th>\n<th><strong>Type / Default</strong></th>\n<th><strong>What It Does</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><code>browser_type</code></strong></td>\n<td><code>\"chromium\"</code>, <code>\"firefox\"</code>, <code>\"webkit\"</code><br><em>(default: <code>\"chromium\"</code>)</em></td>\n<td>Which browser engine to use. <code>\"chromium\"</code> is typical for many sites, <code>\"firefox\"</code> or <code>\"webkit\"</code> for specialized tests.</td>\n</tr>\n<tr>\n<td><strong><code>headless</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>Headless means no visible UI. <code>False</code> is handy for debugging.</td>\n</tr>\n<tr>\n<td><strong><code>viewport_width</code></strong></td>\n<td><code>int</code> (default: <code>1080</code>)</td>\n<td>Initial page width (in px). Useful for testing responsive layouts.</td>\n</tr>\n<tr>\n<td><strong><code>viewport_height</code></strong></td>\n<td><code>int</code> (default: <code>600</code>)</td>\n<td>Initial page height (in px).</td>\n</tr>\n<tr>\n<td><strong><code>proxy</code></strong></td>\n<td><code>str</code> (deprecated)</td>\n<td>Deprecated. Use <code>proxy_config</code> instead. If set, it will be auto-converted internally.</td>\n</tr>\n<tr>\n<td><strong><code>proxy_config</code></strong></td>\n<td><code>dict</code> (default: <code>None</code>)</td>\n<td>For advanced or multi-proxy needs, specify details like <code>{\"server\": \"...\", \"username\": \"...\", ...}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>use_persistent_context</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, uses a <strong>persistent</strong> browser context (keep cookies, sessions across runs). Also sets <code>use_managed_browser=True</code>.</td>\n</tr>\n<tr>\n<td><strong><code>user_data_dir</code></strong></td>\n<td><code>str or None</code> (default: <code>None</code>)</td>\n<td>Directory to store user data (profiles, cookies). Must be set if you want permanent sessions.</td>\n</tr>\n<tr>\n<td><strong><code>ignore_https_errors</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>If <code>True</code>, continues despite invalid certificates (common in dev/staging).</td>\n</tr>\n<tr>\n<td><strong><code>java_script_enabled</code></strong></td>\n<td><code>bool</code> (default: <code>True</code>)</td>\n<td>Disable if you want no JS overhead, or if only static content is needed.</td>\n</tr>\n<tr>\n<td><strong><code>cookies</code></strong></td>\n<td><code>list</code> (default: <code>[]</code>)</td>\n<td>Pre-set cookies, each a dict like <code>{\"name\": \"session\", \"value\": \"...\", \"url\": \"...\"}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>headers</code></strong></td>\n<td><code>dict</code> (default: <code>{}</code>)</td>\n<td>Extra HTTP headers for every request, e.g. <code>{\"Accept-Language\": \"en-US\"}</code>.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent</code></strong></td>\n<td><code>str</code> (default: Chrome-based UA)</td>\n<td>Your custom or random user agent. <code>user_agent_mode=\"random\"</code> can shuffle it.</td>\n</tr>\n<tr>\n<td><strong><code>light_mode</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>Disables some background features for performance gains.</td>\n</tr>\n<tr>\n<td><strong><code>text_mode</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>If <code>True</code>, tries to disable images/other heavy content for speed.</td>\n</tr>\n<tr>\n<td><strong><code>use_managed_browser</code></strong></td>\n<td><code>bool</code> (default: <code>False</code>)</td>\n<td>For advanced “managed” interactions (debugging, CDP usage). Typically set automatically if persistent context is on.</td>\n</tr>\n<tr>\n<td><strong><code>extra_args</code></strong></td>\n<td><code>list</code> (default: <code>[]</code>)</td>\n<td>Additional flags for the underlying browser process, e.g. <code>[\"--disable-extensions\"]</code>.</td>\n</tr>\n<tr>\n<td>- Set <code>headless=False</code> to visually <strong>debug</strong> how pages load or how interactions proceed.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- If you need <strong>authentication</strong> storage or repeated sessions, consider <code>use_persistent_context=True</code> and specify <code>user_data_dir</code>.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- For large pages, you might need a bigger <code>viewport_width</code> and <code>viewport_height</code> to handle dynamic content.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td># 2. <strong>CrawlerRunConfig</strong> – Controlling Each Crawl</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>While <code>BrowserConfig</code> sets up the <strong>environment</strong>, <code>CrawlerRunConfig</code> details <strong>how</strong> each <strong>crawl operation</strong> should behave: caching, content filtering, link or domain blocking, timeouts, JavaScript code, etc.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\nrun_cfg = CrawlerRunConfig(\n    wait_for=<span class=\"hljs-string\">\"css:.main-content\"</span>,\n    word_count_threshold=<span class=\"hljs-number\">15</span>,\n    excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,\n    stream=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable streaming for arun_many()</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>## 2.1 Parameter Highlights</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>### A) <strong>Content Processing</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>------------------------------</td>\n<td>--------------------------------------</td>\n<td>-------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>word_count_threshold</code></strong></td>\n<td><code>int</code> (default: ~200)</td>\n<td>Skips text blocks below X words. Helps ignore trivial sections.</td>\n</tr>\n<tr>\n<td><strong><code>extraction_strategy</code></strong></td>\n<td><code>ExtractionStrategy</code> (default: None)</td>\n<td>If set, extracts structured data (CSS-based, LLM-based, etc.).</td>\n</tr>\n<tr>\n<td><strong><code>markdown_generator</code></strong></td>\n<td><code>MarkdownGenerationStrategy</code> (None)</td>\n<td>If you want specialized markdown output (citations, filtering, chunking, etc.). Can be customized with options such as <code>content_source</code> parameter to select the HTML input source ('cleaned_html', 'raw_html', or 'fit_html').</td>\n</tr>\n<tr>\n<td><strong><code>css_selector</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>Retains only the part of the page matching this selector. Affects the entire extraction process.</td>\n</tr>\n<tr>\n<td><strong><code>target_elements</code></strong></td>\n<td><code>List[str]</code> (None)</td>\n<td>List of CSS selectors for elements to focus on for markdown generation and data extraction, while still processing the entire page for links, media, etc. Provides more flexibility than <code>css_selector</code>.</td>\n</tr>\n<tr>\n<td><strong><code>excluded_tags</code></strong></td>\n<td><code>list</code> (None)</td>\n<td>Removes entire tags (e.g. <code>[\"script\", \"style\"]</code>).</td>\n</tr>\n<tr>\n<td><strong><code>excluded_selector</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>Like <code>css_selector</code> but to exclude. E.g. <code>\"#ads, .tracker\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>only_text</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, tries to extract text-only content.</td>\n</tr>\n<tr>\n<td><strong><code>prettiify</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, beautifies final HTML (slower, purely cosmetic).</td>\n</tr>\n<tr>\n<td><strong><code>keep_data_attributes</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, preserve <code>data-*</code> attributes in cleaned HTML.</td>\n</tr>\n<tr>\n<td><strong><code>remove_forms</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, remove all <code>&lt;form&gt;</code> elements.</td>\n</tr>\n<tr>\n<td>### B) <strong>Caching &amp; Session</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>-------------------------</td>\n<td>------------------------</td>\n<td>------------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>cache_mode</code></strong></td>\n<td><code>CacheMode or None</code></td>\n<td>Controls how caching is handled (<code>ENABLED</code>, <code>BYPASS</code>, <code>DISABLED</code>, etc.). If <code>None</code>, typically defaults to <code>ENABLED</code>.</td>\n</tr>\n<tr>\n<td><strong><code>session_id</code></strong></td>\n<td><code>str or None</code></td>\n<td>Assign a unique ID to reuse a single browser session across multiple <code>arun()</code> calls.</td>\n</tr>\n<tr>\n<td><strong><code>bypass_cache</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, acts like <code>CacheMode.BYPASS</code>.</td>\n</tr>\n<tr>\n<td><strong><code>disable_cache</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, acts like <code>CacheMode.DISABLED</code>.</td>\n</tr>\n<tr>\n<td><strong><code>no_cache_read</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, acts like <code>CacheMode.WRITE_ONLY</code> (writes cache but never reads).</td>\n</tr>\n<tr>\n<td><strong><code>no_cache_write</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, acts like <code>CacheMode.READ_ONLY</code> (reads cache but never writes).</td>\n</tr>\n<tr>\n<td>### C) <strong>Page Navigation &amp; Timing</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>----------------------------</td>\n<td>-------------------------</td>\n<td>----------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>wait_until</code></strong></td>\n<td><code>str</code> (domcontentloaded)</td>\n<td>Condition for navigation to “complete”. Often <code>\"networkidle\"</code> or <code>\"domcontentloaded\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>page_timeout</code></strong></td>\n<td><code>int</code> (60000 ms)</td>\n<td>Timeout for page navigation or JS steps. Increase for slow sites.</td>\n</tr>\n<tr>\n<td><strong><code>wait_for</code></strong></td>\n<td><code>str or None</code></td>\n<td>Wait for a CSS (<code>\"css:selector\"</code>) or JS (<code>\"js:() =&gt; bool\"</code>) condition before content extraction.</td>\n</tr>\n<tr>\n<td><strong><code>wait_for_images</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Wait for images to load before finishing. Slows down if you only want text.</td>\n</tr>\n<tr>\n<td><strong><code>delay_before_return_html</code></strong></td>\n<td><code>float</code> (0.1)</td>\n<td>Additional pause (seconds) before final HTML is captured. Good for last-second updates.</td>\n</tr>\n<tr>\n<td><strong><code>check_robots_txt</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Whether to check and respect robots.txt rules before crawling. If True, caches robots.txt for efficiency.</td>\n</tr>\n<tr>\n<td><strong><code>mean_delay</code></strong> and <strong><code>max_range</code></strong></td>\n<td><code>float</code> (0.1, 0.3)</td>\n<td>If you call <code>arun_many()</code>, these define random delay intervals between crawls, helping avoid detection or rate limits.</td>\n</tr>\n<tr>\n<td><strong><code>semaphore_count</code></strong></td>\n<td><code>int</code> (5)</td>\n<td>Max concurrency for <code>arun_many()</code>. Increase if you have resources for parallel crawls.</td>\n</tr>\n<tr>\n<td>### D) <strong>Page Interaction</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>----------------------------</td>\n<td>--------------------------------</td>\n<td>-----------------------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>js_code</code></strong></td>\n<td><code>str or list[str]</code> (None)</td>\n<td>JavaScript to run <strong>after</strong> <code>wait_for</code> and <code>delay_before_return_html</code>, on the fully-loaded page. E.g. <code>\"document.querySelector('button')?.click();\"</code>.</td>\n</tr>\n<tr>\n<td><strong><code>js_code_before_wait</code></strong></td>\n<td><code>str or list[str]</code> (None)</td>\n<td>JavaScript to run <strong>before</strong> <code>wait_for</code>. Use for triggering loading that <code>wait_for</code> then checks.</td>\n</tr>\n<tr>\n<td><strong><code>js_only</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, indicates we're reusing an existing session and only applying JS. No full reload.</td>\n</tr>\n<tr>\n<td><strong><code>ignore_body_visibility</code></strong></td>\n<td><code>bool</code> (True)</td>\n<td>Skip checking if <code>&lt;body&gt;</code> is visible. Usually best to keep <code>True</code>.</td>\n</tr>\n<tr>\n<td><strong><code>scan_full_page</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, auto-scroll the page to load dynamic content (infinite scroll).</td>\n</tr>\n<tr>\n<td><strong><code>scroll_delay</code></strong></td>\n<td><code>float</code> (0.2)</td>\n<td>Delay between scroll steps when scanning the full page (<code>scan_full_page=True</code>) or capturing full-page screenshots.</td>\n</tr>\n<tr>\n<td><strong><code>process_iframes</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Inlines iframe content for single-page extraction.</td>\n</tr>\n<tr>\n<td><strong><code>flatten_shadow_dom</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Flattens Shadow DOM content into the light DOM before HTML capture. Resolves slots, strips shadow-scoped styles, and force-opens closed shadow roots. Essential for sites built with Web Components.</td>\n</tr>\n<tr>\n<td><strong><code>remove_overlay_elements</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes potential modals/popups blocking the main content.</td>\n</tr>\n<tr>\n<td><strong><code>remove_consent_popups</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes GDPR/cookie consent popups from known CMP providers (OneTrust, Cookiebot, TrustArc, Quantcast, Didomi, Sourcepoint, FundingChoices, etc.). Tries clicking \"Accept All\" first, then falls back to DOM removal.</td>\n</tr>\n<tr>\n<td><strong><code>simulate_user</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Simulate user interactions (mouse movements) to avoid bot detection.</td>\n</tr>\n<tr>\n<td><strong><code>override_navigator</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Override <code>navigator</code> properties in JS for stealth.</td>\n</tr>\n<tr>\n<td><strong><code>magic</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Automatic handling of popups/consent banners. Experimental.</td>\n</tr>\n<tr>\n<td><strong><code>adjust_viewport_to_content</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Resizes viewport to match page content height.</td>\n</tr>\n<tr>\n<td>If your page is a single-page app with repeated JS updates, set <code>js_only=True</code> in subsequent calls, plus a <code>session_id</code> for reusing the same tab.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>### E) <strong>Media Handling</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>--------------------------------------------</td>\n<td>---------------------</td>\n<td>-----------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>screenshot</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Capture a screenshot (base64) in <code>result.screenshot</code>.</td>\n</tr>\n<tr>\n<td><strong><code>screenshot_wait_for</code></strong></td>\n<td><code>float or None</code></td>\n<td>Extra wait time before the screenshot.</td>\n</tr>\n<tr>\n<td><strong><code>screenshot_height_threshold</code></strong></td>\n<td><code>int</code> (~20000)</td>\n<td>If the page is taller than this, alternate screenshot strategies are used.</td>\n</tr>\n<tr>\n<td><strong><code>pdf</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, returns a PDF in <code>result.pdf</code>.</td>\n</tr>\n<tr>\n<td><strong><code>capture_mhtml</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, captures an MHTML snapshot of the page in <code>result.mhtml</code>. MHTML includes all page resources (CSS, images, etc.) in a single file.</td>\n</tr>\n<tr>\n<td><strong><code>image_description_min_word_threshold</code></strong></td>\n<td><code>int</code> (~50)</td>\n<td>Minimum words for an image’s alt text or description to be considered valid.</td>\n</tr>\n<tr>\n<td><strong><code>image_score_threshold</code></strong></td>\n<td><code>int</code> (~3)</td>\n<td>Filter out low-scoring images. The crawler scores images by relevance (size, context, etc.).</td>\n</tr>\n<tr>\n<td><strong><code>exclude_external_images</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Exclude images from other domains.</td>\n</tr>\n<tr>\n<td>### F) <strong>Link/Domain Handling</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>------------------------------</td>\n<td>-------------------------</td>\n<td>-----------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>exclude_social_media_domains</code></strong></td>\n<td><code>list</code> (e.g. Facebook/Twitter)</td>\n<td>A default list can be extended. Any link to these domains is removed from final output.</td>\n</tr>\n<tr>\n<td><strong><code>exclude_external_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Removes all links pointing outside the current domain.</td>\n</tr>\n<tr>\n<td><strong><code>exclude_social_media_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Strips links specifically to social sites (like Facebook or Twitter).</td>\n</tr>\n<tr>\n<td><strong><code>exclude_domains</code></strong></td>\n<td><code>list</code> ([])</td>\n<td>Provide a custom list of domains to exclude (like <code>[\"ads.com\", \"trackers.io\"]</code>).</td>\n</tr>\n<tr>\n<td><strong><code>preserve_https_for_internal_links</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>If <code>True</code>, preserves HTTPS scheme for internal links even when the server redirects to HTTP. Useful for security-conscious crawling.</td>\n</tr>\n<tr>\n<td>### G) <strong>Debug &amp; Logging</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>----------------</td>\n<td>--------------------</td>\n<td>---------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>verbose</code></strong></td>\n<td><code>bool</code> (True)</td>\n<td>Prints logs detailing each step of crawling, interactions, or errors.</td>\n</tr>\n<tr>\n<td><strong><code>log_console</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>Logs the page’s JavaScript console output if you want deeper JS debugging.</td>\n</tr>\n<tr>\n<td>### H) <strong>Virtual Scroll Configuration</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>------------------------------</td>\n<td>------------------------------</td>\n<td>-------------------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>virtual_scroll_config</code></strong></td>\n<td><code>VirtualScrollConfig or dict</code> (None)</td>\n<td>Configuration for handling virtualized scrolling on sites like Twitter/Instagram where content is replaced rather than appended.</td>\n</tr>\n<tr>\n<td>When sites use virtual scrolling (content replaced as you scroll), use <code>VirtualScrollConfig</code>:</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai import VirtualScrollConfig\n\nvirtual_config = VirtualScrollConfig(\n    container_selector=<span class=\"hljs-string\">\"#timeline\"</span>,    <span class=\"hljs-comment\"># CSS selector for scrollable container</span>\n    scroll_count=30,                   <span class=\"hljs-comment\"># Number of times to scroll</span>\n    scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,      <span class=\"hljs-comment\"># How much to scroll: \"container_height\", \"page_height\", or pixels (e.g. 500)</span>\n    wait_after_scroll=0.5             <span class=\"hljs-comment\"># Seconds to wait after each scroll for content to load</span>\n)\n\nconfig = CrawlerRunConfig(\n    virtual_scroll_config=virtual_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>VirtualScrollConfig Parameters:</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>------------------------</td>\n<td>---------------------------</td>\n<td>-------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>container_selector</code></strong></td>\n<td><code>str</code> (required)</td>\n<td>CSS selector for the scrollable container (e.g., <code>\"#feed\"</code>, <code>\".timeline\"</code>)</td>\n</tr>\n<tr>\n<td><strong><code>scroll_count</code></strong></td>\n<td><code>int</code> (10)</td>\n<td>Maximum number of scrolls to perform</td>\n</tr>\n<tr>\n<td><strong><code>scroll_by</code></strong></td>\n<td><code>str or int</code> (\"container_height\")</td>\n<td>Scroll amount: <code>\"container_height\"</code>, <code>\"page_height\"</code>, or pixels (e.g., <code>500</code>)</td>\n</tr>\n<tr>\n<td><strong><code>wait_after_scroll</code></strong></td>\n<td><code>float</code> (0.5)</td>\n<td>Time in seconds to wait after each scroll for new content to load</td>\n</tr>\n<tr>\n<td>- Use <code>virtual_scroll_config</code> when content is <strong>replaced</strong> during scroll (Twitter, Instagram)</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- Use <code>scan_full_page</code> when content is <strong>appended</strong> during scroll (traditional infinite scroll)</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>### I) <strong>URL Matching Configuration</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>------------------------</td>\n<td>------------------------------</td>\n<td>-------------------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>url_matcher</code></strong></td>\n<td><code>UrlMatcher</code> (None)</td>\n<td>Pattern(s) to match URLs against. Can be: string (glob), function, or list of mixed types. <strong>None means match ALL URLs</strong></td>\n</tr>\n<tr>\n<td><strong><code>match_mode</code></strong></td>\n<td><code>MatchMode</code> (MatchMode.OR)</td>\n<td>How to combine multiple matchers in a list: <code>MatchMode.OR</code> (any match) or <code>MatchMode.AND</code> (all must match)</td>\n</tr>\n<tr>\n<td>The <code>url_matcher</code> parameter enables URL-specific configurations when used with <code>arun_many()</code>:</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig, MatchMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.processors.pdf <span class=\"hljs-keyword\">import</span> PDFContentScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.extraction_strategy <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-comment\"># Simple string pattern (glob-style)</span>\npdf_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-string\">\"*.pdf\"</span>,\n    scraping_strategy=PDFContentScrapingStrategy()\n)\n\n<span class=\"hljs-comment\"># Multiple patterns with OR logic (default)</span>\nblog_config = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"*/blog/*\"</span>, <span class=\"hljs-string\">\"*/article/*\"</span>, <span class=\"hljs-string\">\"*/news/*\"</span>],\n    match_mode=MatchMode.OR  <span class=\"hljs-comment\"># Any pattern matches</span>\n)\n\n<span class=\"hljs-comment\"># Function matcher</span>\napi_config = CrawlerRunConfig(\n    url_matcher=<span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'api'</span> <span class=\"hljs-keyword\">in</span> url <span class=\"hljs-keyword\">or</span> url.endswith(<span class=\"hljs-string\">'.json'</span>),\n    <span class=\"hljs-comment\"># Other settings like extraction_strategy</span>\n)\n\n<span class=\"hljs-comment\"># Mixed: String + Function with AND logic</span>\ncomplex_config = CrawlerRunConfig(\n    url_matcher=[\n        <span class=\"hljs-keyword\">lambda</span> url: url.startswith(<span class=\"hljs-string\">'https://'</span>),  <span class=\"hljs-comment\"># Must be HTTPS</span>\n        <span class=\"hljs-string\">\"*.org/*\"</span>,                               <span class=\"hljs-comment\"># Must be .org domain</span>\n        <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'docs'</span> <span class=\"hljs-keyword\">in</span> url                <span class=\"hljs-comment\"># Must contain 'docs'</span>\n    ],\n    match_mode=MatchMode.AND  <span class=\"hljs-comment\"># ALL conditions must match</span>\n)\n\n<span class=\"hljs-comment\"># Combined patterns and functions with AND logic</span>\nsecure_docs = CrawlerRunConfig(\n    url_matcher=[<span class=\"hljs-string\">\"https://*\"</span>, <span class=\"hljs-keyword\">lambda</span> url: <span class=\"hljs-string\">'.doc'</span> <span class=\"hljs-keyword\">in</span> url],\n    match_mode=MatchMode.AND  <span class=\"hljs-comment\"># Must be HTTPS AND contain .doc</span>\n)\n\n<span class=\"hljs-comment\"># Default config - matches ALL URLs</span>\ndefault_config = CrawlerRunConfig()  <span class=\"hljs-comment\"># No url_matcher = matches everything</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>UrlMatcher Types:</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>None (default)</strong>: When <code>url_matcher</code> is None or not set, the config matches ALL URLs</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>String patterns</strong>: Glob-style patterns like <code>\"*.pdf\"</code>, <code>\"*/api/*\"</code>, <code>\"https://*.example.com/*\"</code></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Functions</strong>: <code>lambda url: bool</code> - Custom logic for complex matching</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Lists</strong>: Mix strings and functions, combined with <code>MatchMode.OR</code> or <code>MatchMode.AND</code></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Important Behavior:</strong></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- When passing a list of configs to <code>arun_many()</code>, URLs are matched against each config's <code>url_matcher</code> in order. First match wins!</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- If no config matches a URL and there's no default config (one without <code>url_matcher</code>), the URL will fail with \"No matching configuration found\"</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Both <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> provide a <code>clone()</code> method to create modified copies:</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Create a base configuration</span>\nbase_config = CrawlerRunConfig(\n    cache_mode=CacheMode.ENABLED,\n    word_count_threshold=200\n)\n\n<span class=\"hljs-comment\"># Create variations using clone()</span>\nstream_config = base_config.clone(stream=True)\nno_cache_config = base_config.clone(\n    cache_mode=CacheMode.BYPASS,\n    stream=True\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>The <code>clone()</code> method is particularly useful when you need slightly different configurations for different use cases, without modifying the original config.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>## 2.3 Example Usage</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Configure the browser</span>\n    browser_cfg = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,\n        viewport_width=<span class=\"hljs-number\">1280</span>,\n        viewport_height=<span class=\"hljs-number\">720</span>,\n        proxy_config=<span class=\"hljs-string\">\"http://user:pass@myproxy:8080\"</span>,\n        text_mode=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># Configure the run</span>\n    run_cfg = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        session_id=<span class=\"hljs-string\">\"my_session\"</span>,\n        css_selector=<span class=\"hljs-string\">\"main.article\"</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"script\"</span>, <span class=\"hljs-string\">\"style\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        wait_for=<span class=\"hljs-string\">\"css:.article-loaded\"</span>,\n        screenshot=<span class=\"hljs-literal\">True</span>,\n        stream=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/news\"</span>,\n            config=run_cfg\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Final cleaned_html length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n            <span class=\"hljs-keyword\">if</span> result.screenshot:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Screenshot captured (base64, length):\"</span>, <span class=\"hljs-built_in\">len</span>(result.screenshot))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>## 2.4 Compliance &amp; Ethics</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>-----------------------</td>\n<td>-------------------------</td>\n<td>----------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>check_robots_txt</code></strong></td>\n<td><code>bool</code> (False)</td>\n<td>When True, checks and respects robots.txt rules before crawling. Uses efficient caching with SQLite backend.</td>\n</tr>\n<tr>\n<td><strong><code>user_agent</code></strong></td>\n<td><code>str</code> (None)</td>\n<td>User agent string to identify your crawler. Used for robots.txt checking when enabled.</td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">run_config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    check_robots_txt<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Enable robots.txt compliance</span>\n    user_agent<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"MyBot/1.0\"</span>  <span class=\"hljs-comment\"># Identify your crawler</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td># 3. <strong>LLMConfig</strong> - Setting up LLM providers</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>1. LLMExtractionStrategy</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>2. LLMContentFilter</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>3. JsonCssExtractionStrategy.generate_schema</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>4. JsonXPathExtractionStrategy.generate_schema</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>## 3.1 Parameters</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Parameter</strong></td>\n<td><strong>Type / Default</strong></td>\n<td><strong>What It Does</strong></td>\n</tr>\n<tr>\n<td>-----------------------</td>\n<td>----------------------------------------</td>\n<td>---------------------------------------------------------------------------------------------------------------------------------------</td>\n</tr>\n<tr>\n<td><strong><code>provider</code></strong></td>\n<td><code>\"ollama/llama3\",\"groq/llama3-70b-8192\",\"groq/llama3-8b-8192\", \"openai/gpt-4o-mini\" ,\"openai/gpt-4o\",\"openai/o1-mini\",\"openai/o1-preview\",\"openai/o3-mini\",\"openai/o3-mini-high\",\"anthropic/claude-3-haiku-20240307\",\"anthropic/claude-3-opus-20240229\",\"anthropic/claude-3-sonnet-20240229\",\"anthropic/claude-3-5-sonnet-20240620\",\"gemini/gemini-pro\",\"gemini/gemini-1.5-pro\",\"gemini/gemini-2.0-flash\",\"gemini/gemini-2.0-flash-exp\",\"gemini/gemini-2.0-flash-lite-preview-02-05\",\"deepseek/deepseek-chat\"</code><br><em>(default: <code>\"openai/gpt-4o-mini\"</code>)</em></td>\n<td>Which LLM provider to use.</td>\n</tr>\n<tr>\n<td><strong><code>api_token</code></strong></td>\n<td>1.Optional. When not provided explicitly, api_token will be read from environment variables based on provider. For example: If a gemini model is passed as provider then,<code>\"GEMINI_API_KEY\"</code> will be read from environment variables  <br> 2. API token of LLM provider <br> eg: <code>api_token = \"gsk_1ClHGGJ7Lpn4WGybR7vNWGdyb3FY7zXEw3SCiy0BAVM9lL8CQv\"</code> <br> 3. Environment variable - use with prefix \"env:\" <br> eg:<code>api_token = \"env: GROQ_API_KEY\"</code></td>\n<td>API token to use for the given provider</td>\n</tr>\n<tr>\n<td><strong><code>base_url</code></strong></td>\n<td>Optional. Custom API endpoint</td>\n<td>If your provider has a custom endpoint</td>\n</tr>\n<tr>\n<td>## 3.2 Example Usage</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">llm_config</span> = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>, api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>## 4. Putting It All Together</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Use</strong> <code>BrowserConfig</code> for <strong>global</strong> browser settings: engine, headless, proxy, user agent.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Use</strong> <code>CrawlerRunConfig</code> for each crawl’s <strong>context</strong>: how to filter content, handle caching, wait for dynamic elements, or run JS.</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Pass</strong> both configs to <code>AsyncWebCrawler</code> (the <code>BrowserConfig</code>) and then to <code>arun()</code> (the <code>CrawlerRunConfig</code>).</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>- <strong>Use</strong> <code>LLMConfig</code> for LLM provider configurations that can be used across all extraction, filtering, schema generation tasks. Can be used in - <code>LLMExtractionStrategy</code>, <code>LLMContentFilter</code>, <code>JsonCssExtractionStrategy.generate_schema</code> &amp; <code>JsonXPathExtractionStrategy.generate_schema</code></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"># <span class=\"hljs-keyword\">Create</span> a modified <span class=\"hljs-keyword\">copy</span> <span class=\"hljs-keyword\">with</span> the clone() <span class=\"hljs-keyword\">method</span>\nstream_cfg <span class=\"hljs-operator\">=</span> run_cfg.clone(\n    stream<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>,\n    cache_mode<span class=\"hljs-operator\">=</span>CacheMode.BYPASS\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h1 id=\"crawling-patterns\">Crawling Patterns</h1>\n<h1 id=\"simple-crawling\">Simple Crawling</h1>\n<h2 id=\"basic-usage\">Basic Usage</h2>\n<p>Set up a simple crawl using <code>BrowserConfig</code> and <code>CrawlerRunConfig</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig()  <span class=\"hljs-comment\"># Default browser configuration</span>\n    run_config = CrawlerRunConfig()   <span class=\"hljs-comment\"># Default crawl run configuration</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown)  <span class=\"hljs-comment\"># Print clean markdown content</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"understanding-the-response\">Understanding the Response</h2>\n<p>The <code>arun()</code> method returns a <code>CrawlResult</code> object with several useful properties. Here's a quick overview (see <a href=\"../api/crawl-result.md\">CrawlResult</a> for complete details):\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    markdown_generator=DefaultMarkdownGenerator(\n        content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.6</span>),\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n    config=config\n)\n\n<span class=\"hljs-comment\"># Different content formats</span>\n<span class=\"hljs-built_in\">print</span>(result.html)         <span class=\"hljs-comment\"># Raw HTML</span>\n<span class=\"hljs-built_in\">print</span>(result.cleaned_html) <span class=\"hljs-comment\"># Cleaned HTML</span>\n<span class=\"hljs-built_in\">print</span>(result.markdown.raw_markdown) <span class=\"hljs-comment\"># Raw markdown from cleaned html</span>\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown) <span class=\"hljs-comment\"># Most relevant content in markdown</span>\n\n<span class=\"hljs-comment\"># Check success status</span>\n<span class=\"hljs-built_in\">print</span>(result.success)      <span class=\"hljs-comment\"># True if crawl succeeded</span>\n<span class=\"hljs-built_in\">print</span>(result.status_code)  <span class=\"hljs-comment\"># HTTP status code (e.g., 200, 404)</span>\n\n<span class=\"hljs-comment\"># Access extracted media and links</span>\n<span class=\"hljs-built_in\">print</span>(result.media)        <span class=\"hljs-comment\"># Dictionary of found media (images, videos, audio)</span>\n<span class=\"hljs-built_in\">print</span>(result.links)        <span class=\"hljs-comment\"># Dictionary of internal and external links</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"adding-basic-options\">Adding Basic Options</h2>\n<p>Customize your crawl using <code>CrawlerRunConfig</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">run_config = CrawlerRunConfig(\n    word_count_threshold=<span class=\"hljs-number\">10</span>,        <span class=\"hljs-comment\"># Minimum words per content block</span>\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,    <span class=\"hljs-comment\"># Remove external links</span>\n    remove_overlay_elements=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Remove popups/modals</span>\n    process_iframes=<span class=\"hljs-literal\">True</span>           <span class=\"hljs-comment\"># Process iframe content</span>\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n    config=run_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"handling-errors\">Handling Errors</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">run_config = CrawlerRunConfig()\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n\n<span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Status code: <span class=\"hljs-subst\">{result.status_code}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"logging-and-debugging\">Logging and Debugging</h2>\n<p>Enable verbose logging in <code>BrowserConfig</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">browser_config = BrowserConfig(verbose=True)\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>(<span class=\"hljs-params\">config=browser_config</span>) <span class=\"hljs-keyword\">as</span> crawler:\n    run_config</span> = CrawlerRunConfig()\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"complete-example\">Complete Example</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig(verbose=<span class=\"hljs-literal\">True</span>)\n    run_config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Content filtering</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">'form'</span>, <span class=\"hljs-string\">'header'</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Content processing</span>\n        process_iframes=<span class=\"hljs-literal\">True</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Cache control</span>\n        cache_mode=CacheMode.ENABLED  <span class=\"hljs-comment\"># Use cache if available</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># Print clean content</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Content:\"</span>, result.markdown[:<span class=\"hljs-number\">500</span>])  <span class=\"hljs-comment\"># First 500 chars</span>\n\n            <span class=\"hljs-comment\"># Process images</span>\n            <span class=\"hljs-keyword\">for</span> image <span class=\"hljs-keyword\">in</span> result.media[<span class=\"hljs-string\">\"images\"</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found image: <span class=\"hljs-subst\">{image[<span class=\"hljs-string\">'src'</span>]}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Process links</span>\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links[<span class=\"hljs-string\">\"internal\"</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal link: <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h1 id=\"content-processing\">Content Processing</h1>\n<h1 id=\"markdown-generation-basics\">Markdown Generation Basics</h1>\n<ol>\n<li>How to configure the <strong>Default Markdown Generator</strong>  </li>\n<li>The difference between raw markdown (<code>result.markdown</code>) and filtered markdown (<code>fit_markdown</code>)  <blockquote>\n<ul>\n<li>You know how to configure <code>CrawlerRunConfig</code>.</li>\n</ul>\n</blockquote>\n</li>\n</ol>\n<h2 id=\"1-quick-example\">1. Quick Example</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        markdown_generator=DefaultMarkdownGenerator()\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown Output:\\n\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown)  <span class=\"hljs-comment\"># The unfiltered markdown from the page</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>CrawlerRunConfig( markdown_generator = DefaultMarkdownGenerator() )</code> instructs Crawl4AI to convert the final HTML into markdown at the end of each crawl.<br>\n- The resulting markdown is accessible via <code>result.markdown</code>.<p></p>\n<h2 id=\"2-how-markdown-generation-works\">2. How Markdown Generation Works</h2>\n<h3 id=\"21-html-to-text-conversion-forked-modified\">2.1 HTML-to-Text Conversion (Forked &amp; Modified)</h3>\n<ul>\n<li>Preserves headings, code blocks, bullet points, etc.  </li>\n<li>Removes extraneous tags (scripts, styles) that don’t add meaningful content.  </li>\n<li>Can optionally generate references for links or skip them altogether.</li>\n</ul>\n<h3 id=\"22-link-citations-references\">2.2 Link Citations &amp; References</h3>\n<p>By default, the generator can convert <code>&lt;a href=\"...\"&gt;</code> elements into <code>[text][1]</code> citations, then place the actual links at the bottom of the document. This is handy for research workflows that demand references in a structured manner.</p>\n<h3 id=\"23-optional-content-filters\">2.3 Optional Content Filters</h3>\n<h2 id=\"3-configuring-the-default-markdown-generator\">3. Configuring the Default Markdown Generator</h2>\n<p>You can tweak the output by passing an <code>options</code> dict to <code>DefaultMarkdownGenerator</code>. For example:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Example: ignore all links, don't escape HTML, and wrap text at 80 characters</span>\n    md_generator = DefaultMarkdownGenerator(\n        options={\n            <span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>,\n            <span class=\"hljs-string\">\"escape_html\"</span>: <span class=\"hljs-literal\">False</span>,\n            <span class=\"hljs-string\">\"body_width\"</span>: <span class=\"hljs-number\">80</span>\n        }\n    )\n\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/docs\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown:\\n\"</span>, result.markdown[:<span class=\"hljs-number\">500</span>])  <span class=\"hljs-comment\"># Just a snippet</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    <span class=\"hljs-keyword\">import</span> asyncio\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nSome commonly used <code>options</code>:\n- <strong><code>ignore_links</code></strong> (bool): Whether to remove all hyperlinks in the final markdown.<br>\n- <strong><code>ignore_images</code></strong> (bool): Remove all <code>![image]()</code> references.<br>\n- <strong><code>escape_html</code></strong> (bool): Turn HTML entities into text (default is often <code>True</code>).<br>\n- <strong><code>body_width</code></strong> (int): Wrap text at N characters. <code>0</code> or <code>None</code> means no wrapping.<br>\n- <strong><code>skip_internal_links</code></strong> (bool): If <code>True</code>, omit <code>#localAnchors</code> or internal links referencing the same page.<br>\n- <strong><code>include_sup_sub</code></strong> (bool): Attempt to handle <code>&lt;sup&gt;</code> / <code>&lt;sub&gt;</code> in a more readable way.<p></p>\n<h2 id=\"4-selecting-the-html-source-for-markdown-generation\">4. Selecting the HTML Source for Markdown Generation</h2>\n<p>The <code>content_source</code> parameter allows you to control which HTML content is used as input for markdown generation. This gives you flexibility in how the HTML is processed before conversion to markdown.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Option 1: Use the raw HTML directly from the webpage (before any processing)</span>\n    raw_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"raw_html\"</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Option 2: Use the cleaned HTML (after scraping strategy processing - default)</span>\n    cleaned_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"cleaned_html\"</span>,  <span class=\"hljs-comment\"># This is the default</span>\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Option 3: Use preprocessed HTML optimized for schema extraction</span>\n    fit_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"fit_html\"</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Use one of the generators in your crawler config</span>\n    config = CrawlerRunConfig(\n        markdown_generator=raw_md_generator  <span class=\"hljs-comment\"># Try each of the generators</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown:\\n\"</span>, result.markdown.raw_markdown[:<span class=\"hljs-number\">500</span>])\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    <span class=\"hljs-keyword\">import</span> asyncio\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"html-source-options\">HTML Source Options</h3>\n<ul>\n<li><strong><code>\"cleaned_html\"</code></strong> (default): Uses the HTML after it has been processed by the scraping strategy. This HTML is typically cleaner and more focused on content, with some boilerplate removed.</li>\n<li><strong><code>\"raw_html\"</code></strong>: Uses the original HTML directly from the webpage, before any cleaning or processing. This preserves more of the original content, but may include navigation bars, ads, footers, and other elements that might not be relevant to the main content.</li>\n<li><strong><code>\"fit_html\"</code></strong>: Uses HTML preprocessed for schema extraction. This HTML is optimized for structured data extraction and may have certain elements simplified or removed.</li>\n</ul>\n<h3 id=\"when-to-use-each-option\">When to Use Each Option</h3>\n<ul>\n<li>Use <strong><code>\"cleaned_html\"</code></strong> (default) for most cases where you want a balance of content preservation and noise removal.</li>\n<li>Use <strong><code>\"raw_html\"</code></strong> when you need to preserve all original content, or when the cleaning process is removing content you actually want to keep.</li>\n<li>Use <strong><code>\"fit_html\"</code></strong> when working with structured data or when you need HTML that's optimized for schema extraction.</li>\n</ul>\n<h2 id=\"5-content-filters\">5. Content Filters</h2>\n<h3 id=\"51-bm25contentfilter\">5.1 BM25ContentFilter</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\nbm25_filter = BM25ContentFilter(\n    user_query=<span class=\"hljs-string\">\"machine learning\"</span>,\n    bm25_threshold=<span class=\"hljs-number\">1.2</span>,\n    language=<span class=\"hljs-string\">\"english\"</span>\n)\n\nmd_generator = DefaultMarkdownGenerator(\n    content_filter=bm25_filter,\n    options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n)\n\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>user_query</code></strong>: The term you want to focus on. BM25 tries to keep only content blocks relevant to that query.<br>\n- <strong><code>bm25_threshold</code></strong>: Raise it to keep fewer blocks; lower it to keep more.<br>\n- <strong><code>use_stemming</code></strong> <em>(default <code>True</code>)</em>: Whether to apply stemming to the query and content.\n- <strong><code>language (str)</code></strong>: Language for stemming (default: 'english').<p></p>\n<h3 id=\"52-pruningcontentfilter\">5.2 PruningContentFilter</h3>\n<p>If you <strong>don’t</strong> have a specific query, or if you just want a robust “junk remover,” use <code>PruningContentFilter</code>. It analyzes text density, link density, HTML structure, and known patterns (like “nav,” “footer”) to systematically prune extraneous or repetitive sections.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">from crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n\nprune_filter = <span class=\"hljs-built_in\">PruningContentFilter</span>(\n    threshold=<span class=\"hljs-number\">0.5</span>,\n    threshold_type=<span class=\"hljs-string\">\"fixed\"</span>,  <span class=\"hljs-meta\"># or <span class=\"hljs-string\">\"dynamic\"</span></span>\n    min_word_threshold=<span class=\"hljs-number\">50</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>threshold</code></strong>: Score boundary. Blocks below this score get removed.<br>\n- <strong><code>threshold_type</code></strong>:<br>\n    - <code>\"fixed\"</code>: Straight comparison (<code>score &gt;= threshold</code> keeps the block).<br>\n    - <code>\"dynamic\"</code>: The filter adjusts threshold in a data-driven manner.<br>\n- <strong><code>min_word_threshold</code></strong>: Discard blocks under N words as likely too short or unhelpful.\n- You want a broad cleanup without a user query.  <p></p>\n<h3 id=\"53-llmcontentfilter\">5.3 LLMContentFilter</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, LLMConfig, DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> LLMContentFilter\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Initialize LLM filter with specific instruction</span>\n    <span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-api-token\"</span>), <span class=\"hljs-comment\">#or use environment variable</span>\n        instruction=<span class=\"hljs-string\">\"\"\"\n        Focus on extracting the core educational content.\n        Include:\n        - Key concepts and explanations\n        - Important code examples\n        - Essential technical details\n        Exclude:\n        - Navigation elements\n        - Sidebars\n        - Footer content\n        Format the output as clean markdown with proper code blocks and headers.\n        \"\"\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">4096</span>,  <span class=\"hljs-comment\"># Adjust based on your needs</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n    md_generator = DefaultMarkdownGenerator(\n        content_filter=<span class=\"hljs-built_in\">filter</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n        <span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)  <span class=\"hljs-comment\"># Filtered markdown content</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong>Chunk Processing</strong>: Handles large documents by processing them in chunks (controlled by <code>chunk_token_threshold</code>)\n- <strong>Parallel Processing</strong>: For better performance, use smaller <code>chunk_token_threshold</code> (e.g., 2048 or 4096) to enable parallel processing of content chunks\n1. <strong>Exact Content Preservation</strong>:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n    instruction=<span class=\"hljs-string\">\"\"\"\n    Extract the main educational content while preserving its original wording and substance completely.\n    1. Maintain the exact language and terminology\n    2. Keep all technical explanations and examples intact\n    3. Preserve the original flow and structure\n    4. Remove only clearly irrelevant elements like navigation menus and ads\n    \"\"\"</span>,\n    chunk_token_threshold=<span class=\"hljs-number\">4096</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n2. <strong>Focused Content Extraction</strong>:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n    instruction=<span class=\"hljs-string\">\"\"\"\n    Focus on extracting specific types of content:\n    - Technical documentation\n    - Code examples\n    - API references\n    Reformat the content into clear, well-structured markdown\n    \"\"\"</span>,\n    chunk_token_threshold=<span class=\"hljs-number\">4096</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<blockquote>\n<p><strong>Performance Tip</strong>: Set a smaller <code>chunk_token_threshold</code> (e.g., 2048 or 4096) to enable parallel processing of content chunks. The default value is infinity, which processes the entire content as a single chunk.</p>\n</blockquote>\n<h2 id=\"6-using-fit-markdown\">6. Using Fit Markdown</h2>\n<p>When a content filter is active, the library produces two forms of markdown inside <code>result.markdown</code>:\n1. <strong><code>raw_markdown</code></strong>: The full unfiltered markdown.<br>\n2. <strong><code>fit_markdown</code></strong>: A “fit” version where the filter has removed or trimmed noisy segments.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        markdown_generator=DefaultMarkdownGenerator(\n            content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.6</span>),\n            options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n        )\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://news.example.com/tech\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw markdown:\\n\"</span>, result.markdown)\n\n            <span class=\"hljs-comment\"># If a filter is used, we also have .fit_markdown:</span>\n            md_object = result.markdown  <span class=\"hljs-comment\"># or your equivalent</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Filtered markdown:\\n\"</span>, md_object.fit_markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"7-the-markdowngenerationresult-object\">7. The <code>MarkdownGenerationResult</code> Object</h2>\n<p>If your library stores detailed markdown output in an object like <code>MarkdownGenerationResult</code>, you’ll see fields such as:\n- <strong><code>raw_markdown</code></strong>: The direct HTML-to-markdown transformation (no filtering).<br>\n- <strong><code>markdown_with_citations</code></strong>: A version that moves links to reference-style footnotes.<br>\n- <strong><code>references_markdown</code></strong>: A separate string or section containing the gathered references.<br>\n- <strong><code>fit_markdown</code></strong>: The filtered markdown if you used a content filter.<br>\n- <strong><code>fit_html</code></strong>: The corresponding HTML snippet used to generate <code>fit_markdown</code> (helpful for debugging or advanced usage).\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-swift\">md_obj <span class=\"hljs-operator\">=</span> result.markdown  # your library’s naming may vary\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"RAW:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.raw_markdown)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"CITED:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.markdown_with_citations)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"REFERENCES:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.references_markdown)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"FIT:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.fit_markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- You can supply <code>raw_markdown</code> to an LLM if you want the entire text.<br>\n- Or feed <code>fit_markdown</code> into a vector database to reduce token usage.<br>\n- <code>references_markdown</code> can help you keep track of link provenance.<p></p>\n<h2 id=\"8-combining-filters-bm25-pruning-in-two-passes\">8. Combining Filters (BM25 + Pruning) in Two Passes</h2>\n<p>You might want to <strong>prune out</strong> noisy boilerplate first (with <code>PruningContentFilter</code>), and then <strong>rank what’s left</strong> against a user query (with <code>BM25ContentFilter</code>). You don’t have to crawl the page twice. Instead:\n1. <strong>First pass</strong>: Apply <code>PruningContentFilter</code> directly to the raw HTML from <code>result.html</code> (the crawler’s downloaded HTML).<br>\n2. <strong>Second pass</strong>: Take the pruned HTML (or text) from step 1, and feed it into <code>BM25ContentFilter</code>, focusing on a user query.</p>\n<h3 id=\"two-pass-example\">Two-Pass Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter, BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> bs4 <span class=\"hljs-keyword\">import</span> BeautifulSoup\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Crawl with minimal or no markdown generator, just get raw HTML</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># If you only want raw HTML, you can skip passing a markdown_generator</span>\n        <span class=\"hljs-comment\"># or provide one but focus on .html in this example</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/tech-article\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success <span class=\"hljs-keyword\">or</span> <span class=\"hljs-keyword\">not</span> result.html:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed or no HTML content.\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        raw_html = result.html\n\n        <span class=\"hljs-comment\"># 2. First pass: PruningContentFilter on raw HTML</span>\n        pruning_filter = PruningContentFilter(threshold=<span class=\"hljs-number\">0.5</span>, min_word_threshold=<span class=\"hljs-number\">50</span>)\n\n        <span class=\"hljs-comment\"># filter_content returns a list of \"text chunks\" or cleaned HTML sections</span>\n        pruned_chunks = pruning_filter.filter_content(raw_html)\n        <span class=\"hljs-comment\"># This list is basically pruned content blocks, presumably in HTML or text form</span>\n\n        <span class=\"hljs-comment\"># For demonstration, let's combine these chunks back into a single HTML-like string</span>\n        <span class=\"hljs-comment\"># or you could do further processing. It's up to your pipeline design.</span>\n        pruned_html = <span class=\"hljs-string\">\"\\n\"</span>.join(pruned_chunks)\n\n        <span class=\"hljs-comment\"># 3. Second pass: BM25ContentFilter with a user query</span>\n        bm25_filter = BM25ContentFilter(\n            user_query=<span class=\"hljs-string\">\"machine learning\"</span>,\n            bm25_threshold=<span class=\"hljs-number\">1.2</span>,\n            language=<span class=\"hljs-string\">\"english\"</span>\n        )\n\n        <span class=\"hljs-comment\"># returns a list of text chunks</span>\n        bm25_chunks = bm25_filter.filter_content(pruned_html)  \n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> bm25_chunks:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Nothing matched the BM25 query after pruning.\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># 4. Combine or display final results</span>\n        final_text = <span class=\"hljs-string\">\"\\n---\\n\"</span>.join(bm25_chunks)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"==== PRUNED OUTPUT (first pass) ====\"</span>)\n        <span class=\"hljs-built_in\">print</span>(pruned_html[:<span class=\"hljs-number\">500</span>], <span class=\"hljs-string\">\"... (truncated)\"</span>)  <span class=\"hljs-comment\"># preview</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n==== BM25 OUTPUT (second pass) ====\"</span>)\n        <span class=\"hljs-built_in\">print</span>(final_text[:<span class=\"hljs-number\">500</span>], <span class=\"hljs-string\">\"... (truncated)\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"whats-happening\">What’s Happening?</h3>\n<p>1. <strong>Raw HTML</strong>: We crawl once and store the raw HTML in <code>result.html</code>.<br>\n4. <strong>BM25ContentFilter</strong>: We feed the pruned string into <code>BM25ContentFilter</code> with a user query. This second pass further narrows the content to chunks relevant to “machine learning.”\n<strong>No Re-Crawling</strong>: We used <code>raw_html</code> from the first pass, so there’s no need to run <code>arun()</code> again—<strong>no second network request</strong>.</p>\n<h3 id=\"tips-variations\">Tips &amp; Variations</h3>\n<ul>\n<li><strong>Plain Text vs. HTML</strong>: If your pruned output is mostly text, BM25 can still handle it; just keep in mind it expects a valid string input. If you supply partial HTML (like <code>\"&lt;p&gt;some text&lt;/p&gt;\"</code>), it will parse it as HTML.  </li>\n<li><strong>Adjust Thresholds</strong>: If you see too much or too little text in step one, tweak <code>threshold=0.5</code> or <code>min_word_threshold=50</code>. Similarly, <code>bm25_threshold=1.2</code> can be raised/lowered for more or fewer chunks in step two.</li>\n</ul>\n<h3 id=\"one-pass-combination\">One-Pass Combination?</h3>\n<h2 id=\"9-common-pitfalls-tips\">9. Common Pitfalls &amp; Tips</h2>\n<p>1. <strong>No Markdown Output?</strong><br>\n2. <strong>Performance Considerations</strong><br>\n   - Very large pages with multiple filters can be slower. Consider <code>cache_mode</code> to avoid re-downloading.<br>\n3. <strong>Take Advantage of <code>fit_markdown</code></strong><br>\n4. <strong>Adjusting <code>html2text</code> Options</strong><br>\n   - If you see lots of raw HTML slipping into the text, turn on <code>escape_html</code>.<br>\n   - If code blocks look messy, experiment with <code>mark_code</code> or <code>handle_code_in_pre</code>.</p>\n<h2 id=\"10-summary-next-steps\">10. Summary &amp; Next Steps</h2>\n<ul>\n<li>Configure the <strong>DefaultMarkdownGenerator</strong> with HTML-to-text options.  </li>\n<li>Select different HTML sources using the <code>content_source</code> parameter.  </li>\n<li>Distinguish between raw and filtered markdown (<code>fit_markdown</code>).  </li>\n<li>Leverage the <code>MarkdownGenerationResult</code> object to handle different forms of output (citations, references, etc.).</li>\n</ul>\n<h1 id=\"fit-markdown-with-pruning-bm25\">Fit Markdown with Pruning &amp; BM25</h1>\n<h2 id=\"1-how-fit-markdown-works\">1. How “Fit Markdown” Works</h2>\n<h3 id=\"11-the-content_filter\">1.1 The <code>content_filter</code></h3>\n<p>In <strong><code>CrawlerRunConfig</code></strong>, you can specify a <strong><code>content_filter</code></strong> to shape how content is pruned or ranked before final markdown generation. A filter’s logic is applied <strong>before</strong> or <strong>during</strong> the HTML→Markdown process, producing:\n- <strong><code>result.markdown.raw_markdown</code></strong> (unfiltered)\n- <strong><code>result.markdown.fit_markdown</code></strong> (filtered or “fit” version)\n- <strong><code>result.markdown.fit_html</code></strong> (the corresponding HTML snippet that produced <code>fit_markdown</code>)</p>\n<h3 id=\"12-common-filters\">1.2 Common Filters</h3>\n<h2 id=\"2-pruningcontentfilter\">2. PruningContentFilter</h2>\n<h3 id=\"21-usage-example\">2.1 Usage Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Step 1: Create a pruning filter</span>\n    prune_filter = PruningContentFilter(\n        <span class=\"hljs-comment\"># Lower → more content retained, higher → more content pruned</span>\n        threshold=<span class=\"hljs-number\">0.45</span>,           \n        <span class=\"hljs-comment\"># \"fixed\" or \"dynamic\"</span>\n        threshold_type=<span class=\"hljs-string\">\"dynamic\"</span>,  \n        <span class=\"hljs-comment\"># Ignore nodes with &lt;5 words</span>\n        min_word_threshold=<span class=\"hljs-number\">5</span>      \n    )\n\n    <span class=\"hljs-comment\"># Step 2: Insert it into a Markdown Generator</span>\n    md_generator = DefaultMarkdownGenerator(content_filter=prune_filter)\n\n    <span class=\"hljs-comment\"># Step 3: Pass it to CrawlerRunConfig</span>\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, \n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># 'fit_markdown' is your pruned content, focusing on \"denser\" text</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.fit_markdown))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"22-key-parameters\">2.2 Key Parameters</h3>\n<ul>\n<li><strong><code>min_word_threshold</code></strong> (int): If a block has fewer words than this, it’s pruned.  </li>\n<li><strong><code>threshold_type</code></strong> (str):</li>\n<li><code>\"fixed\"</code> → each node must exceed <code>threshold</code> (0–1).  </li>\n<li><code>\"dynamic\"</code> → node scoring adjusts according to tag type, text/link density, etc.  </li>\n<li><strong><code>threshold</code></strong> (float, default ~0.48): The base or “anchor” cutoff.  </li>\n<li><strong>Link density</strong> – Penalizes sections that are mostly links.  </li>\n<li><strong>Tag importance</strong> – e.g., an <code>&lt;article&gt;</code> or <code>&lt;p&gt;</code> might be more important than a <code>&lt;div&gt;</code>.  </li>\n</ul>\n<h2 id=\"3-bm25contentfilter\">3. BM25ContentFilter</h2>\n<h3 id=\"31-usage-example\">3.1 Usage Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1) A BM25 filter with a user query</span>\n    bm25_filter = BM25ContentFilter(\n        user_query=<span class=\"hljs-string\">\"startup fundraising tips\"</span>,\n        <span class=\"hljs-comment\"># Adjust for stricter or looser results</span>\n        bm25_threshold=<span class=\"hljs-number\">1.2</span>  \n    )\n\n    <span class=\"hljs-comment\"># 2) Insert into a Markdown Generator</span>\n    md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)\n\n    <span class=\"hljs-comment\"># 3) Pass to crawler config</span>\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, \n            config=config\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown (BM25 query-based):\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-parameters\">3.2 Parameters</h3>\n<ul>\n<li><strong><code>user_query</code></strong> (str, optional): E.g. <code>\"machine learning\"</code>. If blank, the filter tries to glean a query from page metadata.  </li>\n<li><strong><code>bm25_threshold</code></strong> (float, default 1.0):  </li>\n<li>Higher → fewer chunks but more relevant.  </li>\n<li>Lower → more inclusive.  <blockquote>\n<p>In more advanced scenarios, you might see parameters like <code>language</code>, <code>case_sensitive</code>, or <code>priority_tags</code> to refine how text is tokenized or weighted.</p>\n</blockquote>\n</li>\n</ul>\n<h2 id=\"4-accessing-the-fit-output\">4. Accessing the “Fit” Output</h2>\n<p>After the crawl, your “fit” content is found in <strong><code>result.markdown.fit_markdown</code></strong>. \n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">fit_md</span> = result.markdown.fit_markdown\n<span class=\"hljs-attr\">fit_html</span> = result.markdown.fit_html\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nIf the content filter is <strong>BM25</strong>, you might see additional logic or references in <code>fit_markdown</code> that highlight relevant segments. If it’s <strong>Pruning</strong>, the text is typically well-cleaned but not necessarily matched to a query.<p></p>\n<h2 id=\"5-code-patterns-recap\">5. Code Patterns Recap</h2>\n<h3 id=\"51-pruning\">5.1 Pruning</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">prune_filter = PruningContentFilter(\n    threshold=0.5,\n    threshold_type=<span class=\"hljs-string\">\"fixed\"</span>,\n    min_word_threshold=10\n)\nmd_generator = DefaultMarkdownGenerator(content_filter=prune_filter)\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"52-bm25\">5.2 BM25</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">bm25_filter = BM25ContentFilter(\n    user_query=<span class=\"hljs-string\">\"health benefits fruit\"</span>,\n    bm25_threshold=1.2\n)\nmd_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"6-combining-with-word_count_threshold-exclusions\">6. Combining with “word_count_threshold” &amp; Exclusions</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">10</span>,\n    excluded_tags<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>, <span class=\"hljs-string\">\"header\"</span><span class=\"hljs-punctuation\">]</span>,\n    exclude_external_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    markdown_generator<span class=\"hljs-punctuation\">=</span>DefaultMarkdownGenerator<span class=\"hljs-punctuation\">(</span>\n        content_filter<span class=\"hljs-punctuation\">=</span>PruningContentFilter<span class=\"hljs-punctuation\">(</span>threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.5</span><span class=\"hljs-punctuation\">)</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. The crawler’s <code>excluded_tags</code> are removed from the HTML first.<br>\n3. The final “fit” content is generated in <code>result.markdown.fit_markdown</code>.<p></p>\n<h2 id=\"7-custom-filters\">7. Custom Filters</h2>\n<p>If you need a different approach (like a specialized ML model or site-specific heuristics), you can create a new class inheriting from <code>RelevantContentFilter</code> and implement <code>filter_content(html)</code>. Then inject it into your <strong>markdown generator</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> RelevantContentFilter\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">MyCustomFilter</span>(<span class=\"hljs-title class_ inherited__\">RelevantContentFilter</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">filter_content</span>(<span class=\"hljs-params\">self, html, min_word_threshold=<span class=\"hljs-literal\">None</span></span>):\n        <span class=\"hljs-comment\"># parse HTML, implement custom logic</span>\n        <span class=\"hljs-keyword\">return</span> [block <span class=\"hljs-keyword\">for</span> block <span class=\"hljs-keyword\">in</span> ... <span class=\"hljs-keyword\">if</span> ... some condition...]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. Subclass <code>RelevantContentFilter</code>.<br>\n2. Implement <code>filter_content(...)</code>.<br>\n3. Use it in your <code>DefaultMarkdownGenerator(content_filter=MyCustomFilter(...))</code>.<p></p>\n<h2 id=\"8-final-thoughts\">8. Final Thoughts</h2>\n<ul>\n<li><strong>Summaries</strong>: Quickly get the important text from a cluttered page.  </li>\n<li><strong>Search</strong>: Combine with <strong>BM25</strong> to produce content relevant to a query.  </li>\n<li><strong>BM25ContentFilter</strong>: Perfect for query-based extraction or searching.  </li>\n<li>Combine with <strong><code>excluded_tags</code>, <code>exclude_external_links</code>, <code>word_count_threshold</code></strong> to refine your final “fit” text.  </li>\n<li>Fit markdown ends up in <strong><code>result.markdown.fit_markdown</code></strong>; eventually <strong><code>result.markdown.fit_markdown</code></strong> in future versions.</li>\n<li>Last Updated: 2025-01-01</li>\n</ul>\n<h1 id=\"content-selection\">Content Selection</h1>\n<p>Crawl4AI provides multiple ways to <strong>select</strong>, <strong>filter</strong>, and <strong>refine</strong> the content from your crawls. Whether you need to target a specific CSS region, exclude entire tags, filter out external links, or remove certain domains and images, <strong><code>CrawlerRunConfig</code></strong> offers a wide range of parameters.</p>\n<h2 id=\"1-css-based-selection\">1. CSS-Based Selection</h2>\n<p>There are two ways to select content from a page: using <code>css_selector</code> or the more flexible <code>target_elements</code>.</p>\n<h3 id=\"11-using-css_selector\">1.1 Using <code>css_selector</code></h3>\n<p>A straightforward way to <strong>limit</strong> your crawl results to a certain region of the page is <strong><code>css_selector</code></strong> in <strong><code>CrawlerRunConfig</code></strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># e.g., first 30 items from Hacker News</span>\n        css_selector=<span class=\"hljs-string\">\".athing:nth-child(-n+30)\"</span>  \n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Partial HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Result</strong>: Only elements matching that selector remain in <code>result.cleaned_html</code>.<p></p>\n<h3 id=\"12-using-target_elements\">1.2 Using <code>target_elements</code></h3>\n<p>The <code>target_elements</code> parameter provides more flexibility by allowing you to target <strong>multiple elements</strong> for content extraction while preserving the entire page context for other features:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Target article body and sidebar, but not other content</span>\n        target_elements=[<span class=\"hljs-string\">\"article.main-content\"</span>, <span class=\"hljs-string\">\"aside.sidebar\"</span>]\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/blog-post\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown focused on target elements\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Links from entire page still available:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Key difference</strong>: With <code>target_elements</code>, the markdown generation and structural data extraction focus on those elements, but other page elements (like links, images, and tables) are still extracted from the entire page. This gives you fine-grained control over what appears in your markdown content while preserving full page context for link analysis and media collection.<p></p>\n<h2 id=\"2-content-filtering-exclusions\">2. Content Filtering &amp; Exclusions</h2>\n<h3 id=\"21-basic-overview\">2.1 Basic Overview</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    <span class=\"hljs-comment\"># Content thresholds</span>\n    word_count_threshold=<span class=\"hljs-number\">10</span>,        <span class=\"hljs-comment\"># Minimum words per block</span>\n\n    <span class=\"hljs-comment\"># Tag exclusions</span>\n    excluded_tags=[<span class=\"hljs-string\">'form'</span>, <span class=\"hljs-string\">'header'</span>, <span class=\"hljs-string\">'footer'</span>, <span class=\"hljs-string\">'nav'</span>],\n\n    <span class=\"hljs-comment\"># Link filtering</span>\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,    \n    exclude_social_media_links=<span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-comment\"># Block entire domains</span>\n    exclude_domains=[<span class=\"hljs-string\">\"adtrackers.com\"</span>, <span class=\"hljs-string\">\"spammynews.org\"</span>],    \n    exclude_social_media_domains=[<span class=\"hljs-string\">\"facebook.com\"</span>, <span class=\"hljs-string\">\"twitter.com\"</span>],\n\n    <span class=\"hljs-comment\"># Media filtering</span>\n    exclude_external_images=<span class=\"hljs-literal\">True</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>word_count_threshold</code></strong>: Ignores text blocks under X words. Helps skip trivial blocks like short nav or disclaimers.<br>\n- <strong><code>excluded_tags</code></strong>: Removes entire tags (<code>&lt;form&gt;</code>, <code>&lt;header&gt;</code>, <code>&lt;footer&gt;</code>, etc.).<br>\n- <strong>Link Filtering</strong>:<br>\n  - <code>exclude_external_links</code>: Strips out external links and may remove them from <code>result.links</code>.<br>\n  - <code>exclude_social_media_links</code>: Removes links pointing to known social media domains.<br>\n  - <code>exclude_domains</code>: A custom list of domains to block if discovered in links.<br>\n  - <code>exclude_social_media_domains</code>: A curated list (override or add to it) for social media sites.<br>\n- <strong>Media Filtering</strong>:<br>\n  - <code>exclude_external_images</code>: Discards images not hosted on the same domain as the main page (or its subdomains).\nBy default in case you set <code>exclude_social_media_links=True</code>, the following social media domains are excluded:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">[\n    <span class=\"hljs-string\">'facebook.com'</span>,\n    <span class=\"hljs-string\">'twitter.com'</span>,\n    <span class=\"hljs-string\">'x.com'</span>,\n    <span class=\"hljs-string\">'linkedin.com'</span>,\n    <span class=\"hljs-string\">'instagram.com'</span>,\n    <span class=\"hljs-string\">'pinterest.com'</span>,\n    <span class=\"hljs-string\">'tiktok.com'</span>,\n    <span class=\"hljs-string\">'snapchat.com'</span>,\n    <span class=\"hljs-string\">'reddit.com'</span>,\n]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"22-example-usage\">2.2 Example Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        css_selector=<span class=\"hljs-string\">\"main.content\"</span>, \n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"ads.com\"</span>, <span class=\"hljs-string\">\"spammytrackers.net\"</span>],\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"3-handling-iframes\">3. Handling Iframes</h2>\n<p>Some sites embed content in <code>&lt;iframe&gt;</code> tags. If you want that inline:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\">config <span class=\"hljs-operator\">=</span> CrawlerRunConfig(\n    # <span class=\"hljs-keyword\">Merge</span> iframe content <span class=\"hljs-keyword\">into</span> the <span class=\"hljs-keyword\">final</span> output\n    process_iframes<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>,    \n    remove_overlay_elements<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        process_iframes=<span class=\"hljs-literal\">True</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.org/iframe-demo\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Iframe-merged length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"31-flattening-shadow-dom\">3.1 Flattening Shadow DOM</h2>\n<p>Sites built with <strong>Web Components</strong> (Stencil, Lit, Shoelace, Angular Elements, etc.) render content inside Shadow DOM — an encapsulated sub-tree invisible to <code>page.content()</code>. Set <code>flatten_shadow_dom=True</code> to extract it:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    flatten_shadow_dom<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    wait_until<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"load\"</span>,\n    delay_before_return_html<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3.0</span>,  <span class=\"hljs-comment\"># give components time to hydrate</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        flatten_shadow_dom=<span class=\"hljs-literal\">True</span>,\n        wait_until=<span class=\"hljs-string\">\"load\"</span>,\n        delay_before_return_html=<span class=\"hljs-number\">3.0</span>,\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://store.boschrexroth.com/en/us/p/hydraulic-cylinder-r900999011\"</span>,\n            config=config,\n        )\n        <span class=\"hljs-comment\"># Without flatten_shadow_dom: ~1 KB markdown (breadcrumbs only)</span>\n        <span class=\"hljs-comment\"># With flatten_shadow_dom:   ~33 KB (product description, specs, downloads)</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nWhen enabled, Crawl4AI also injects an init script that force-opens closed shadow roots. The flattener resolves <code>&lt;slot&gt;</code> projections and strips shadow-scoped <code>&lt;style&gt;</code> tags, producing clean HTML for the downstream scraping/markdown pipeline.<p></p>\n<p><strong>Execution order</strong>: <code>flatten_shadow_dom</code> runs right before HTML capture, after all waits and JS execution:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">js_code_before_wait → wait_for → delay → js_code → flatten_shadow_dom → page capture\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>For a full runnable example, see <a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/shadow_dom_crawling.py\"><code>shadow_dom_crawling.py</code></a>.</p>\n<h2 id=\"4-structured-extraction-examples\">4. Structured Extraction Examples</h2>\n<h3 id=\"41-pattern-based-with-jsoncssextractionstrategy\">4.1 Pattern-Based with <code>JsonCssExtractionStrategy</code></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Minimal schema for repeated items</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"News Items\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"tr.athing\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.titleline a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.titleline a\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, \n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>\n            }\n        ]\n    }\n\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Content filtering</span>\n        excluded_tags=[<span class=\"hljs-string\">\"form\"</span>, <span class=\"hljs-string\">\"header\"</span>],\n        exclude_domains=[<span class=\"hljs-string\">\"adsite.com\"</span>],\n\n        <span class=\"hljs-comment\"># CSS selection or entire page</span>\n        css_selector=<span class=\"hljs-string\">\"table.itemlist\"</span>,\n\n        <span class=\"hljs-comment\"># No caching for demonstration</span>\n        cache_mode=CacheMode.BYPASS,\n\n        <span class=\"hljs-comment\"># Extraction strategy</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema)\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>, \n            config=config\n        )\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Sample extracted item:\"</span>, data[:<span class=\"hljs-number\">1</span>])  <span class=\"hljs-comment\"># Show first item</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"42-llm-based-extraction\">4.2 LLM-Based Extraction</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">ArticleData</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    headline: <span class=\"hljs-built_in\">str</span>\n    summary: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    llm_strategy = LLMExtractionStrategy(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4\"</span>,api_token=<span class=\"hljs-string\">\"sk-YOUR_API_KEY\"</span>)\n        schema=ArticleData.schema(),\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract 'headline' and a short 'summary' from the content.\"</span>\n    )\n\n    config = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        word_count_threshold=<span class=\"hljs-number\">20</span>,\n        extraction_strategy=llm_strategy\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n        article = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(article)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- Filters out external links (<code>exclude_external_links=True</code>).<br>\n- Ignores very short text blocks (<code>word_count_threshold=20</code>).<br>\n- Passes the final HTML to your LLM strategy for an AI-driven parse.<p></p>\n<h2 id=\"5-comprehensive-example\">5. Comprehensive Example</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_main_articles</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span></span>):\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"ArticleBlock\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.article-block\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"headline\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"summary\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".summary\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"metadata\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"nested\"</span>,\n                <span class=\"hljs-string\">\"fields\"</span>: [\n                    {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"author\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".author\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n                    {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"date\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".date\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>}\n                ]\n            }\n        ]\n    }\n\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Keep only #main-content</span>\n        css_selector=<span class=\"hljs-string\">\"#main-content\"</span>,\n\n        <span class=\"hljs-comment\"># Filtering</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],  \n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"somebadsite.com\"</span>],\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Extraction</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=config)\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">None</span>\n        <span class=\"hljs-keyword\">return</span> json.loads(result.extracted_content)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    articles = <span class=\"hljs-keyword\">await</span> extract_main_articles(<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>)\n    <span class=\"hljs-keyword\">if</span> articles:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted Articles:\"</span>, articles[:<span class=\"hljs-number\">2</span>])  <span class=\"hljs-comment\"># Show first 2</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong>CSS</strong> scoping with <code>#main-content</code>.<br>\n- Multiple <strong>exclude_</strong> parameters to remove domains, external images, etc.<br>\n- A <strong>JsonCssExtractionStrategy</strong> to parse repeated article blocks.<p></p>\n<h2 id=\"6-scraping-modes\">6. Scraping Modes</h2>\n<p>Crawl4AI uses <code>LXMLWebScrapingStrategy</code> (LXML-based) as the default scraping strategy for HTML content processing. This strategy offers excellent performance, especially for large HTML documents.\n<strong>Note:</strong> For backward compatibility, <code>WebScrapingStrategy</code> is still available as an alias for <code>LXMLWebScrapingStrategy</code>.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LXMLWebScrapingStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Default configuration already uses LXMLWebScrapingStrategy</span>\n    config = CrawlerRunConfig()\n\n    <span class=\"hljs-comment\"># Or explicitly specify it if desired</span>\n    config_explicit = CrawlerRunConfig(\n        scraping_strategy=LXMLWebScrapingStrategy()\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>, \n            config=config\n        )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nYou can also create your own custom scraping strategy by inheriting from <code>ContentScrapingStrategy</code>. The strategy must return a <code>ScrapingResult</code> object with the following structure:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> ContentScrapingStrategy, ScrapingResult, MediaItem, Media, Link, Links\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CustomScrapingStrategy</span>(<span class=\"hljs-title class_ inherited__\">ContentScrapingStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">scrap</span>(<span class=\"hljs-params\">self, url: <span class=\"hljs-built_in\">str</span>, html: <span class=\"hljs-built_in\">str</span>, **kwargs</span>) -&gt; ScrapingResult:\n        <span class=\"hljs-comment\"># Implement your custom scraping logic here</span>\n        <span class=\"hljs-keyword\">return</span> ScrapingResult(\n            cleaned_html=<span class=\"hljs-string\">\"&lt;html&gt;...&lt;/html&gt;\"</span>,  <span class=\"hljs-comment\"># Cleaned HTML content</span>\n            success=<span class=\"hljs-literal\">True</span>,                     <span class=\"hljs-comment\"># Whether scraping was successful</span>\n            media=Media(\n                images=[                      <span class=\"hljs-comment\"># List of images found</span>\n                    MediaItem(\n                        src=<span class=\"hljs-string\">\"https://example.com/image.jpg\"</span>,\n                        alt=<span class=\"hljs-string\">\"Image description\"</span>,\n                        desc=<span class=\"hljs-string\">\"Surrounding text\"</span>,\n                        score=<span class=\"hljs-number\">1</span>,\n                        <span class=\"hljs-built_in\">type</span>=<span class=\"hljs-string\">\"image\"</span>,\n                        group_id=<span class=\"hljs-number\">1</span>,\n                        <span class=\"hljs-built_in\">format</span>=<span class=\"hljs-string\">\"jpg\"</span>,\n                        width=<span class=\"hljs-number\">800</span>\n                    )\n                ],\n                videos=[],                    <span class=\"hljs-comment\"># List of videos (same structure as images)</span>\n                audios=[]                     <span class=\"hljs-comment\"># List of audio files (same structure as images)</span>\n            ),\n            links=Links(\n                internal=[                    <span class=\"hljs-comment\"># List of internal links</span>\n                    Link(\n                        href=<span class=\"hljs-string\">\"https://example.com/page\"</span>,\n                        text=<span class=\"hljs-string\">\"Link text\"</span>,\n                        title=<span class=\"hljs-string\">\"Link title\"</span>,\n                        base_domain=<span class=\"hljs-string\">\"example.com\"</span>\n                    )\n                ],\n                external=[]                   <span class=\"hljs-comment\"># List of external links (same structure)</span>\n            ),\n            metadata={                        <span class=\"hljs-comment\"># Additional metadata</span>\n                <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"Page Title\"</span>,\n                <span class=\"hljs-string\">\"description\"</span>: <span class=\"hljs-string\">\"Page description\"</span>\n            }\n        )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">ascrap</span>(<span class=\"hljs-params\">self, url: <span class=\"hljs-built_in\">str</span>, html: <span class=\"hljs-built_in\">str</span>, **kwargs</span>) -&gt; ScrapingResult:\n        <span class=\"hljs-comment\"># For simple cases, you can use the sync version</span>\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-keyword\">await</span> asyncio.to_thread(self.scrap, url, html, **kwargs)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"performance-considerations\">Performance Considerations</h3>\n<ul>\n<li>Fast processing of large HTML documents (especially &gt;100KB)</li>\n<li>Efficient memory usage</li>\n<li>Good handling of well-formed HTML</li>\n<li>Robust table detection and extraction</li>\n</ul>\n<h3 id=\"backward-compatibility\">Backward Compatibility</h3>\n<p>For users upgrading from earlier versions:\n- <code>WebScrapingStrategy</code> is now an alias for <code>LXMLWebScrapingStrategy</code>\n- Existing code using <code>WebScrapingStrategy</code> will continue to work without modification\n- No changes are required to your existing code</p>\n<h2 id=\"7-combining-css-selection-methods\">7. Combining CSS Selection Methods</h2>\n<p>You can combine <code>css_selector</code> and <code>target_elements</code> in powerful ways to achieve fine-grained control over your output:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Target specific content but preserve page context</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Focus markdown on main content and sidebar</span>\n        target_elements=[<span class=\"hljs-string\">\"#main-content\"</span>, <span class=\"hljs-string\">\".sidebar\"</span>],\n\n        <span class=\"hljs-comment\"># Global filters applied to entire page</span>\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>, <span class=\"hljs-string\">\"header\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Use basic content thresholds</span>\n        word_count_threshold=<span class=\"hljs-number\">15</span>,\n\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/article\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Content focuses on specific elements, but all links still analyzed\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal links: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">'internal'</span>, []))}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"External links: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">'external'</span>, []))}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- Links, images and other page data still give you the full context of the page\n- Content filtering still applies globally<p></p>\n<h2 id=\"8-conclusion\">8. Conclusion</h2>\n<p>By mixing <strong>target_elements</strong> or <strong>css_selector</strong> scoping, <strong>content filtering</strong> parameters, and advanced <strong>extraction strategies</strong>, you can precisely <strong>choose</strong> which data to keep. Key parameters in <strong><code>CrawlerRunConfig</code></strong> for content selection include:\n1. <strong><code>target_elements</code></strong> – Array of CSS selectors to focus markdown generation and data extraction, while preserving full page context for links and media.\n2. <strong><code>css_selector</code></strong> – Basic scoping to an element or region for all extraction processes.<br>\n3. <strong><code>word_count_threshold</code></strong> – Skip short blocks.<br>\n4. <strong><code>excluded_tags</code></strong> – Remove entire HTML tags.<br>\n5. <strong><code>exclude_external_links</code></strong>, <strong><code>exclude_social_media_links</code></strong>, <strong><code>exclude_domains</code></strong> – Filter out unwanted links or domains.<br>\n6. <strong><code>exclude_external_images</code></strong> – Remove images from external sources.<br>\n7. <strong><code>process_iframes</code></strong> – Merge iframe content if needed.  </p>\n<h1 id=\"page-interaction\">Page Interaction</h1>\n<ol>\n<li>Click “Load More” buttons  </li>\n<li>Fill forms and submit them  </li>\n<li>Wait for elements or data to appear  </li>\n<li>Reuse sessions across multiple steps  </li>\n</ol>\n<h2 id=\"1-javascript-execution\">1. JavaScript Execution</h2>\n<h3 id=\"basic-execution\">Basic Execution</h3>\n<p><strong><code>js_code</code></strong> in <strong><code>CrawlerRunConfig</code></strong> accepts either a single JS string or a list of JS snippets.<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Single JS command</span>\n    config = CrawlerRunConfig(\n        js_code=<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># Example site</span>\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawled length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n    <span class=\"hljs-comment\"># Multiple commands</span>\n    js_commands = [\n        <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        <span class=\"hljs-comment\"># 'More' link on Hacker News</span>\n        <span class=\"hljs-string\">\"document.querySelector('a.morelink')?.click();\"</span>,  \n    ]\n    config = CrawlerRunConfig(js_code=js_commands)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># Another pass</span>\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"After scroll+click, length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Relevant <code>CrawlerRunConfig</code> params</strong>:\n- <strong><code>js_code</code></strong>: A string or list of strings with JavaScript to run after the page loads.\n- <strong><code>js_only</code></strong>: If set to <code>True</code> on subsequent calls, indicates we’re continuing an existing session without a new full navigation.<br>\n- <strong><code>session_id</code></strong>: If you want to keep the same page across multiple calls, specify an ID.<p></p>\n<h2 id=\"2-wait-conditions\">2. Wait Conditions</h2>\n<h3 id=\"21-css-based-waiting\">2.1 CSS-Based Waiting</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Wait for at least 30 items on Hacker News</span>\n        wait_for=<span class=\"hljs-string\">\"css:.athing:nth-child(30)\"</span>  \n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"We have at least 30 items loaded!\"</span>)\n        <span class=\"hljs-comment\"># Rough check</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Total items in HTML:\"</span>, result.cleaned_html.count(<span class=\"hljs-string\">\"athing\"</span>))  \n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>wait_for=\"css:...\"</code></strong>: Tells the crawler to wait until that CSS selector is present.<p></p>\n<h3 id=\"22-javascript-based-waiting\">2.2 JavaScript-Based Waiting</h3>\n<p>For more complex conditions (e.g., waiting for content length to exceed a threshold), prefix <code>js:</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">wait_condition</span> = <span class=\"hljs-string\">\"\"\"() =&gt; {\n    const items = document.querySelectorAll('.athing');\n    return items.length &gt; 50;  // Wait for at least 51 items\n}\"\"\"</span>\n\n<span class=\"hljs-attr\">config</span> = CrawlerRunConfig(wait_for=f<span class=\"hljs-string\">\"js:{wait_condition}\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<strong>Behind the Scenes</strong>: Crawl4AI keeps polling the JS function until it returns <code>true</code> or a timeout occurs.<p></p>\n<h2 id=\"3-handling-dynamic-content\">3. Handling Dynamic Content</h2>\n<h3 id=\"31-load-more-example-hacker-news-more-link\">3.1 Load More Example (Hacker News “More” Link)</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Step 1: Load initial Hacker News page</span>\n    config = CrawlerRunConfig(\n        wait_for=<span class=\"hljs-string\">\"css:.athing:nth-child(30)\"</span>  <span class=\"hljs-comment\"># Wait for 30 items</span>\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Initial items loaded.\"</span>)\n\n        <span class=\"hljs-comment\"># Step 2: Let's scroll and click the \"More\" link</span>\n        load_more_js = [\n            <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n            <span class=\"hljs-comment\"># The \"More\" link at page bottom</span>\n            <span class=\"hljs-string\">\"document.querySelector('a.morelink')?.click();\"</span>  \n        ]\n\n        next_page_conf = CrawlerRunConfig(\n            js_code=load_more_js,\n            wait_for=<span class=\"hljs-string\">\"\"\"js:() =&gt; {\n                return document.querySelectorAll('.athing').length &gt; 30;\n            }\"\"\"</span>,\n            <span class=\"hljs-comment\"># Mark that we do not re-navigate, but run JS in the same session:</span>\n            js_only=<span class=\"hljs-literal\">True</span>,\n            session_id=<span class=\"hljs-string\">\"hn_session\"</span>\n        )\n\n        <span class=\"hljs-comment\"># Re-use the same crawler session</span>\n        result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># same URL but continuing session</span>\n            config=next_page_conf\n        )\n        total_items = result2.cleaned_html.count(<span class=\"hljs-string\">\"athing\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Items after load-more:\"</span>, total_items)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>session_id=\"hn_session\"</code></strong>: Keep the same page across multiple calls to <code>arun()</code>.\n- <strong><code>js_only=True</code></strong>: We’re not performing a full reload, just applying JS in the existing page.\n- <strong><code>wait_for</code></strong> with <code>js:</code>: Wait for item count to grow beyond 30.<p></p>\n<h3 id=\"32-form-interaction\">3.2 Form Interaction</h3>\n<p>If the site has a search or login form, you can fill fields and submit them with <strong><code>js_code</code></strong>. For instance, if GitHub had a local search form:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">js_form_interaction = <span class=\"hljs-string\">\"\"\"\ndocument.querySelector('#your-search').value = 'TypeScript commits';\ndocument.querySelector('form').submit();\n\"\"\"</span>\n\nconfig = CrawlerRunConfig(\n    js_code=js_form_interaction,\n    wait_for=<span class=\"hljs-string\">\"css:.commit\"</span>\n)\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://github.com/search\"</span>, config=config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"4-timing-control\">4. Timing Control</h2>\n<p>1. <strong><code>page_timeout</code></strong> (ms): Overall page load or script execution time limit.<br>\n2. <strong><code>delay_before_return_html</code></strong> (seconds): Wait an extra moment before capturing the final HTML.<br>\n3. <strong><code>mean_delay</code></strong> &amp; <strong><code>max_range</code></strong>: If you call <code>arun_many()</code> with multiple URLs, these add a random pause between each request.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    page_timeout=60000,  <span class=\"hljs-comment\"># 60s limit</span>\n    delay_before_return_html=2.5\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"5-multi-step-interaction-example\">5. Multi-Step Interaction Example</h2>\n<p>Below is a simplified script that does multiple “Load More” clicks on GitHub’s TypeScript commits page. It <strong>re-uses</strong> the same session to accumulate new commits each time. The code includes the relevant <strong><code>CrawlerRunConfig</code></strong> parameters you’d rely on.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">multi_page_commits</span>():\n    browser_cfg = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Visible for demonstration</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n    session_id = <span class=\"hljs-string\">\"github_ts_commits\"</span>\n\n    base_wait = <span class=\"hljs-string\">\"\"\"js:() =&gt; {\n        const commits = document.querySelectorAll('li.Box-sc-g0xbh4-0 h4');\n        return commits.length &gt; 0;\n    }\"\"\"</span>\n\n    <span class=\"hljs-comment\"># Step 1: Load initial commits</span>\n    config1 = CrawlerRunConfig(\n        wait_for=base_wait,\n        session_id=session_id,\n        cache_mode=CacheMode.BYPASS,\n        <span class=\"hljs-comment\"># Not using js_only yet since it's our first load</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>,\n            config=config1\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Initial commits loaded. Count:\"</span>, result.cleaned_html.count(<span class=\"hljs-string\">\"commit\"</span>))\n\n        <span class=\"hljs-comment\"># Step 2: For subsequent pages, we run JS to click 'Next Page' if it exists</span>\n        js_next_page = <span class=\"hljs-string\">\"\"\"\n        const selector = 'a[data-testid=\"pagination-next-button\"]';\n        const button = document.querySelector(selector);\n        if (button) button.click();\n        \"\"\"</span>\n\n        <span class=\"hljs-comment\"># Wait until new commits appear</span>\n        wait_for_more = <span class=\"hljs-string\">\"\"\"js:() =&gt; {\n            const commits = document.querySelectorAll('li.Box-sc-g0xbh4-0 h4');\n            if (!window.firstCommit &amp;&amp; commits.length&gt;0) {\n                window.firstCommit = commits[0].textContent;\n                return false;\n            }\n            // If top commit changes, we have new commits\n            const topNow = commits[0]?.textContent.trim();\n            return topNow &amp;&amp; topNow !== window.firstCommit;\n        }\"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">2</span>):  <span class=\"hljs-comment\"># let's do 2 more \"Next\" pages</span>\n            config_next = CrawlerRunConfig(\n                session_id=session_id,\n                js_code=js_next_page,\n                wait_for=wait_for_more,\n                js_only=<span class=\"hljs-literal\">True</span>,       <span class=\"hljs-comment\"># We're continuing from the open tab</span>\n                cache_mode=CacheMode.BYPASS\n            )\n            result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(\n                url=<span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>,\n                config=config_next\n            )\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page+<span class=\"hljs-number\">2</span>}</span> commits count:\"</span>, result2.cleaned_html.count(<span class=\"hljs-string\">\"commit\"</span>))\n\n        <span class=\"hljs-comment\"># Optionally kill session</span>\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">await</span> multi_page_commits()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>session_id</code></strong>: Keep the same page open.<br>\n- <strong><code>js_code</code></strong> + <strong><code>wait_for</code></strong> + <strong><code>js_only=True</code></strong>: We do partial refreshes, waiting for new commits to appear.<br>\n- <strong><code>cache_mode=CacheMode.BYPASS</code></strong> ensures we always see fresh data each step.<p></p>\n<h2 id=\"6-combine-interaction-with-extraction\">6. Combine Interaction with Extraction</h2>\n<p>Once dynamic content is loaded, you can attach an <strong><code>extraction_strategy</code></strong> (like <code>JsonCssExtractionStrategy</code> or <code>LLMExtractionStrategy</code>). For example:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">from crawl4ai import JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Commits\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"li.Box-sc-g0xbh4-0\"</span>,\n    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h4.markdown-title\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span>\n    <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    session_id<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"ts_commits_session\"</span>,\n    js_code<span class=\"hljs-punctuation\">=</span>js_next_page,\n    wait_for<span class=\"hljs-punctuation\">=</span>wait_for_more,\n    extraction_strategy<span class=\"hljs-punctuation\">=</span>JsonCssExtractionStrategy<span class=\"hljs-punctuation\">(</span><span class=\"hljs-keyword\">schema</span><span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nWhen done, check <code>result.extracted_content</code> for the JSON.<p></p>\n<h2 id=\"7-relevant-crawlerrunconfig-parameters\">7. Relevant <code>CrawlerRunConfig</code> Parameters</h2>\n<p>Below are the key interaction-related parameters in <code>CrawlerRunConfig</code>. For a full list, see <a href=\"../api/parameters.md\">Configuration Parameters</a>.\n- <strong><code>js_code</code></strong>: JavaScript to run after initial load.<br>\n- <strong><code>js_only</code></strong>: If <code>True</code>, no new page navigation—only JS in the existing session.<br>\n- <strong><code>wait_for</code></strong>: CSS (<code>\"css:...\"</code>) or JS (<code>\"js:...\"</code>) expression to wait for.<br>\n- <strong><code>session_id</code></strong>: Reuse the same page across calls.<br>\n- <strong><code>cache_mode</code></strong>: Whether to read/write from the cache or bypass.<br>\n- <strong><code>remove_overlay_elements</code></strong>: Remove certain popups automatically.\n- <strong><code>remove_consent_popups</code></strong>: Remove GDPR/cookie consent popups from known CMP providers (OneTrust, Cookiebot, Didomi, etc.).\n- <strong><code>simulate_user</code>, <code>override_navigator</code>, <code>magic</code></strong>: Anti-bot or \"human-like\" interactions.</p>\n<h2 id=\"8-conclusion_1\">8. Conclusion</h2>\n<p>1. <strong>Execute JavaScript</strong> for scrolling, clicks, or form filling.<br>\n2. <strong>Wait</strong> for CSS or custom JS conditions before capturing data.<br>\n4. Combine with <strong>structured extraction</strong> for dynamic sites.</p>\n<h2 id=\"9-virtual-scrolling\">9. Virtual Scrolling</h2>\n<p>For sites that use <strong>virtual scrolling</strong> (where content is replaced rather than appended as you scroll, like Twitter or Instagram), Crawl4AI provides a dedicated <code>VirtualScrollConfig</code>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, VirtualScrollConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_twitter_timeline</span>():\n    <span class=\"hljs-comment\"># Configure virtual scroll for Twitter-like feeds</span>\n    virtual_config = VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\"[data-testid='primaryColumn']\"</span>,  <span class=\"hljs-comment\"># Twitter's main column</span>\n        scroll_count=<span class=\"hljs-number\">30</span>,                <span class=\"hljs-comment\"># Scroll 30 times</span>\n        scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,   <span class=\"hljs-comment\"># Scroll by container height each time</span>\n        wait_after_scroll=<span class=\"hljs-number\">1.0</span>          <span class=\"hljs-comment\"># Wait 1 second after each scroll</span>\n    )\n\n    config = CrawlerRunConfig(\n        virtual_scroll_config=virtual_config\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://twitter.com/search?q=AI\"</span>,\n            config=config\n        )\n        <span class=\"hljs-comment\"># result.html now contains ALL tweets from the virtual scroll</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"virtual-scroll-vs-javascript-scrolling\">Virtual Scroll vs JavaScript Scrolling</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th>Virtual Scroll</th>\n<th>JS Code Scrolling</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Use Case</strong></td>\n<td>Content replaced during scroll</td>\n<td>Content appended or simple scroll</td>\n</tr>\n<tr>\n<td><strong>Configuration</strong></td>\n<td><code>VirtualScrollConfig</code> object</td>\n<td><code>js_code</code> with scroll commands</td>\n</tr>\n<tr>\n<td><strong>Automatic Merging</strong></td>\n<td>Yes - merges all unique content</td>\n<td>No - captures final state only</td>\n</tr>\n<tr>\n<td><strong>Best For</strong></td>\n<td>Twitter, Instagram, virtual tables</td>\n<td>Traditional pages, load more buttons</td>\n</tr>\n</tbody>\n</table>\n<h1 id=\"link-media\">Link &amp; Media</h1>\n<ol>\n<li>Extract links (internal, external) from crawled pages  </li>\n<li>Filter or exclude specific domains (e.g., social media or custom domains)  </li>\n<li>Access and ma### 3.2 Excluding Images\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_all_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li>You don't need image data in your results</li>\n<li>You're crawling image-heavy pages that cause memory issues</li>\n<li>You want to focus only on text content</li>\n<li>Configure your crawler to exclude or prioritize certain images\nBelow is a revised version of the <strong>Link Extraction</strong> and <strong>Media Extraction</strong> sections that includes example data structures showing how links and media items are stored in <code>CrawlResult</code>. Feel free to adjust any field names or descriptions to match your actual output.</li>\n</ol>\n<h2 id=\"1-link-extraction\">1. Link Extraction</h2>\n<h3 id=\"11-resultlinks\">1.1 <code>result.links</code></h3>\n<p>When you call <code>arun()</code> or <code>arun_many()</code> on a URL, Crawl4AI automatically extracts links and stores them in the <code>links</code> field of <code>CrawlResult</code>. By default, the crawler tries to distinguish <strong>internal</strong> links (same domain) from <strong>external</strong> links (different domains).\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://www.example.com\"</span>)\n    <span class=\"hljs-keyword\">if</span> result.success:\n        internal_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n        external_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> internal links.\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> external links.\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.media)}</span> media items.\"</span>)\n\n        <span class=\"hljs-comment\"># Each link is typically a dictionary with fields like:</span>\n        <span class=\"hljs-comment\"># { \"href\": \"...\", \"text\": \"...\", \"title\": \"...\", \"base_domain\": \"...\" }</span>\n        <span class=\"hljs-keyword\">if</span> internal_links:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Sample Internal Link:\"</span>, internal_links[<span class=\"hljs-number\">0</span>])\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">result.links = {\n  <span class=\"hljs-string\">\"internal\"</span>: [\n    {\n      <span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"https://kidocode.com/\"</span>,\n      <span class=\"hljs-string\">\"text\"</span>: <span class=\"hljs-string\">\"\"</span>,\n      <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"\"</span>,\n      <span class=\"hljs-string\">\"base_domain\"</span>: <span class=\"hljs-string\">\"kidocode.com\"</span>\n    },\n    {\n      <span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"https://kidocode.com/degrees/technology\"</span>,\n      <span class=\"hljs-string\">\"text\"</span>: <span class=\"hljs-string\">\"Technology Degree\"</span>,\n      <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"KidoCode Tech Program\"</span>,\n      <span class=\"hljs-string\">\"base_domain\"</span>: <span class=\"hljs-string\">\"kidocode.com\"</span>\n    },\n    <span class=\"hljs-comment\"># ...</span>\n  ],\n  <span class=\"hljs-string\">\"external\"</span>: [\n    <span class=\"hljs-comment\"># possibly other links leading to third-party sites</span>\n  ]\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>href</code></strong>: The raw hyperlink URL.<br>\n- <strong><code>text</code></strong>: The link text (if any) within the <code>&lt;a&gt;</code> tag.<br>\n- <strong><code>title</code></strong>: The <code>title</code> attribute of the link (if present).<br>\n- <strong><code>base_domain</code></strong>: The domain extracted from <code>href</code>. Helpful for filtering or grouping by domain.<p></p>\n<h2 id=\"2-advanced-link-head-extraction-scoring\">2. Advanced Link Head Extraction &amp; Scoring</h2>\n<p>Ever wanted to not just extract links, but also get the actual content (title, description, metadata) from those linked pages? And score them for relevance? This is exactly what Link Head Extraction does - it fetches the <code>&lt;head&gt;</code> section from each discovered link and scores them using multiple algorithms.</p>\n<h3 id=\"21-why-link-head-extraction\">2.1 Why Link Head Extraction?</h3>\n<ol>\n<li><strong>Fetching head content</strong> from each link (title, description, meta tags)</li>\n<li><strong>Combining scores intelligently</strong> to give you a final relevance ranking</li>\n</ol>\n<h3 id=\"22-complete-working-example\">2.2 Complete Working Example</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LinkPreviewConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_link_heads_example</span>():\n    <span class=\"hljs-string\">\"\"\"\n    Complete example showing link head extraction with scoring.\n    This will crawl a documentation site and extract head content from internal links.\n    \"\"\"</span>\n\n    <span class=\"hljs-comment\"># Configure link head extraction</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Enable link head extraction with detailed configuration</span>\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Extract from internal links</span>\n            include_external=<span class=\"hljs-literal\">False</span>,          <span class=\"hljs-comment\"># Skip external links for this example</span>\n            max_links=<span class=\"hljs-number\">10</span>,                   <span class=\"hljs-comment\"># Limit to 10 links for demo</span>\n            concurrency=<span class=\"hljs-number\">5</span>,                  <span class=\"hljs-comment\"># Process 5 links simultaneously</span>\n            timeout=<span class=\"hljs-number\">10</span>,                     <span class=\"hljs-comment\"># 10 second timeout per link</span>\n            query=<span class=\"hljs-string\">\"API documentation guide\"</span>, <span class=\"hljs-comment\"># Query for contextual scoring</span>\n            score_threshold=<span class=\"hljs-number\">0.3</span>,            <span class=\"hljs-comment\"># Only include links scoring above 0.3</span>\n            verbose=<span class=\"hljs-literal\">True</span>                    <span class=\"hljs-comment\"># Show detailed progress</span>\n        ),\n        <span class=\"hljs-comment\"># Enable intrinsic scoring (URL quality, text relevance)</span>\n        score_links=<span class=\"hljs-literal\">True</span>,\n        <span class=\"hljs-comment\"># Keep output clean</span>\n        only_text=<span class=\"hljs-literal\">True</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Crawl a documentation site (great for testing)</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✅ Successfully crawled: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📄 Page title: <span class=\"hljs-subst\">{result.metadata.get(<span class=\"hljs-string\">'title'</span>, <span class=\"hljs-string\">'No title'</span>)}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Access links (now enhanced with head data and scores)</span>\n            internal_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n            external_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n🔗 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> internal links\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🌍 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(external_links)}</span> external links\"</span>)\n\n            <span class=\"hljs-comment\"># Count links with head data</span>\n            links_with_head = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> internal_links \n                             <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"head_data\"</span>) <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>]\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🧠 Links with head data extracted: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(links_with_head)}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Show the top 3 scoring links</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n🏆 Top 3 Links with Full Scoring:\"</span>)\n            <span class=\"hljs-keyword\">for</span> i, link <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(links_with_head[:<span class=\"hljs-number\">3</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>. <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Link Text: '<span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'text'</span>, <span class=\"hljs-string\">'No text'</span>)[:<span class=\"hljs-number\">50</span>]}</span>...'\"</span>)\n\n                <span class=\"hljs-comment\"># Show all three score types</span>\n                intrinsic = link.get(<span class=\"hljs-string\">'intrinsic_score'</span>)\n                contextual = link.get(<span class=\"hljs-string\">'contextual_score'</span>) \n                total = link.get(<span class=\"hljs-string\">'total_score'</span>)\n\n                <span class=\"hljs-keyword\">if</span> intrinsic <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📊 Intrinsic Score: <span class=\"hljs-subst\">{intrinsic:<span class=\"hljs-number\">.2</span>f}</span>/10.0 (URL quality &amp; context)\"</span>)\n                <span class=\"hljs-keyword\">if</span> contextual <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   🎯 Contextual Score: <span class=\"hljs-subst\">{contextual:<span class=\"hljs-number\">.3</span>f}</span> (BM25 relevance to query)\"</span>)\n                <span class=\"hljs-keyword\">if</span> total <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   ⭐ Total Score: <span class=\"hljs-subst\">{total:<span class=\"hljs-number\">.3</span>f}</span> (combined final score)\"</span>)\n\n                <span class=\"hljs-comment\"># Show extracted head data</span>\n                head_data = link.get(<span class=\"hljs-string\">\"head_data\"</span>, {})\n                <span class=\"hljs-keyword\">if</span> head_data:\n                    title = head_data.get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"No title\"</span>)\n                    description = head_data.get(<span class=\"hljs-string\">\"meta\"</span>, {}).get(<span class=\"hljs-string\">\"description\"</span>, <span class=\"hljs-string\">\"No description\"</span>)\n\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📰 Title: <span class=\"hljs-subst\">{title[:<span class=\"hljs-number\">60</span>]}</span>...\"</span>)\n                    <span class=\"hljs-keyword\">if</span> description:\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📝 Description: <span class=\"hljs-subst\">{description[:<span class=\"hljs-number\">80</span>]}</span>...\"</span>)\n\n                    <span class=\"hljs-comment\"># Show extraction status</span>\n                    status = link.get(<span class=\"hljs-string\">\"head_extraction_status\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   ✅ Extraction Status: <span class=\"hljs-subst\">{status}</span>\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"❌ Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-comment\"># Run the example</span>\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(extract_link_heads_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-yaml\"><span class=\"hljs-string\">✅</span> <span class=\"hljs-attr\">Successfully crawled:</span> <span class=\"hljs-string\">https://docs.python.org/3/</span>\n<span class=\"hljs-string\">📄</span> <span class=\"hljs-attr\">Page title:</span> <span class=\"hljs-number\">3.13</span><span class=\"hljs-number\">.5</span> <span class=\"hljs-string\">Documentation</span>\n<span class=\"hljs-string\">🔗</span> <span class=\"hljs-string\">Found</span> <span class=\"hljs-number\">53</span> <span class=\"hljs-string\">internal</span> <span class=\"hljs-string\">links</span>\n<span class=\"hljs-string\">🌍</span> <span class=\"hljs-string\">Found</span> <span class=\"hljs-number\">1</span> <span class=\"hljs-string\">external</span> <span class=\"hljs-string\">links</span>\n<span class=\"hljs-string\">🧠</span> <span class=\"hljs-attr\">Links with head data extracted:</span> <span class=\"hljs-number\">10</span>\n\n<span class=\"hljs-string\">🏆</span> <span class=\"hljs-attr\">Top 3 Links with Full Scoring:</span>\n\n<span class=\"hljs-number\">1</span><span class=\"hljs-string\">.</span> <span class=\"hljs-string\">https://docs.python.org/3.15/</span>\n   <span class=\"hljs-attr\">Link Text:</span> <span class=\"hljs-string\">'Python 3.15 (in development)...'</span>\n   <span class=\"hljs-string\">📊</span> <span class=\"hljs-attr\">Intrinsic Score:</span> <span class=\"hljs-number\">4.17</span><span class=\"hljs-string\">/10.0</span> <span class=\"hljs-string\">(URL</span> <span class=\"hljs-string\">quality</span> <span class=\"hljs-string\">&amp;</span> <span class=\"hljs-string\">context)</span>\n   <span class=\"hljs-string\">🎯</span> <span class=\"hljs-attr\">Contextual Score:</span> <span class=\"hljs-number\">1.000</span> <span class=\"hljs-string\">(BM25</span> <span class=\"hljs-string\">relevance</span> <span class=\"hljs-string\">to</span> <span class=\"hljs-string\">query)</span>\n   <span class=\"hljs-string\">⭐</span> <span class=\"hljs-attr\">Total Score:</span> <span class=\"hljs-number\">5.917</span> <span class=\"hljs-string\">(combined</span> <span class=\"hljs-string\">final</span> <span class=\"hljs-string\">score)</span>\n   <span class=\"hljs-string\">📰</span> <span class=\"hljs-attr\">Title:</span> <span class=\"hljs-number\">3.15</span><span class=\"hljs-string\">.0a0</span> <span class=\"hljs-string\">Documentation...</span>\n   <span class=\"hljs-string\">📝</span> <span class=\"hljs-attr\">Description:</span> <span class=\"hljs-string\">The</span> <span class=\"hljs-string\">official</span> <span class=\"hljs-string\">Python</span> <span class=\"hljs-string\">documentation...</span>\n   <span class=\"hljs-string\">✅</span> <span class=\"hljs-attr\">Extraction Status:</span> <span class=\"hljs-string\">valid</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"23-configuration-deep-dive\">2.3 Configuration Deep Dive</h3>\n<p>The <code>LinkPreviewConfig</code> class supports these options:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LinkPreviewConfig\n\nlink_preview_config = LinkPreviewConfig(\n    <span class=\"hljs-comment\"># BASIC SETTINGS</span>\n    verbose=<span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Show detailed logs (recommended for learning)</span>\n\n    <span class=\"hljs-comment\"># LINK FILTERING</span>\n    include_internal=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Include same-domain links</span>\n    include_external=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Include different-domain links</span>\n    max_links=<span class=\"hljs-number\">50</span>,                   <span class=\"hljs-comment\"># Maximum links to process (prevents overload)</span>\n\n    <span class=\"hljs-comment\"># PATTERN FILTERING</span>\n    include_patterns=[               <span class=\"hljs-comment\"># Only process links matching these patterns</span>\n        <span class=\"hljs-string\">\"*/docs/*\"</span>, \n        <span class=\"hljs-string\">\"*/api/*\"</span>, \n        <span class=\"hljs-string\">\"*/reference/*\"</span>\n    ],\n    exclude_patterns=[               <span class=\"hljs-comment\"># Skip links matching these patterns</span>\n        <span class=\"hljs-string\">\"*/login*\"</span>,\n        <span class=\"hljs-string\">\"*/admin*\"</span>\n    ],\n\n    <span class=\"hljs-comment\"># PERFORMANCE SETTINGS</span>\n    concurrency=<span class=\"hljs-number\">10</span>,                  <span class=\"hljs-comment\"># How many links to process simultaneously</span>\n    timeout=<span class=\"hljs-number\">5</span>,                      <span class=\"hljs-comment\"># Seconds to wait per link</span>\n\n    <span class=\"hljs-comment\"># RELEVANCE SCORING</span>\n    query=<span class=\"hljs-string\">\"machine learning API\"</span>,    <span class=\"hljs-comment\"># Query for BM25 contextual scoring</span>\n    score_threshold=<span class=\"hljs-number\">0.3</span>,            <span class=\"hljs-comment\"># Only include links above this score</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"24-understanding-the-three-score-types\">2.4 Understanding the Three Score Types</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># High intrinsic score indicators:</span>\n<span class=\"hljs-comment\"># ✅ Clean URL structure (docs.python.org/api/reference)</span>\n<span class=\"hljs-comment\"># ✅ Meaningful link text (\"API Reference Guide\")</span>\n<span class=\"hljs-comment\"># ✅ Relevant to page context</span>\n<span class=\"hljs-comment\"># ✅ Not buried deep in navigation</span>\n\n<span class=\"hljs-comment\"># Low intrinsic score indicators:</span>\n<span class=\"hljs-comment\"># ❌ Random URLs (site.com/x7f9g2h)</span>\n<span class=\"hljs-comment\"># ❌ No link text or generic text (\"Click here\")</span>\n<span class=\"hljs-comment\"># ❌ Unrelated to page content</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nOnly available when you provide a <code>query</code>. Uses BM25 algorithm against head content:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-shell\"><span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">Example: query = <span class=\"hljs-string\">\"machine learning tutorial\"</span></span>\n<span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">High contextual score: Link to <span class=\"hljs-string\">\"Complete Machine Learning Guide\"</span></span>\n<span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">Low contextual score: Link to <span class=\"hljs-string\">\"Privacy Policy\"</span></span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># When both scores available: (intrinsic * 0.3) + (contextual * 0.7)</span>\n<span class=\"hljs-comment\"># When only intrinsic: uses intrinsic score</span>\n<span class=\"hljs-comment\"># When only contextual: uses contextual score</span>\n<span class=\"hljs-comment\"># When neither: not calculated</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"25-practical-use-cases\">2.5 Practical Use Cases</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">research_assistant</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            include_external=<span class=\"hljs-literal\">True</span>,\n            include_patterns=[<span class=\"hljs-string\">\"*/docs/*\"</span>, <span class=\"hljs-string\">\"*/tutorial/*\"</span>, <span class=\"hljs-string\">\"*/guide/*\"</span>],\n            query=<span class=\"hljs-string\">\"machine learning neural networks\"</span>,\n            max_links=<span class=\"hljs-number\">20</span>,\n            score_threshold=<span class=\"hljs-number\">0.5</span>,  <span class=\"hljs-comment\"># Only high-relevance links</span>\n            verbose=<span class=\"hljs-literal\">True</span>\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://scikit-learn.org/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># Get high-scoring links</span>\n            good_links = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n                         <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"total_score\"</span>, <span class=\"hljs-number\">0</span>) &gt; <span class=\"hljs-number\">0.7</span>]\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🎯 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(good_links)}</span> highly relevant links:\"</span>)\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> good_links[:<span class=\"hljs-number\">5</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"⭐ <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'total_score'</span>]:<span class=\"hljs-number\">.3</span>f}</span> - <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   <span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'head_data'</span>, {}</span>).get('title', 'No title')}\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">api_discovery</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            include_patterns=[<span class=\"hljs-string\">\"*/api/*\"</span>, <span class=\"hljs-string\">\"*/reference/*\"</span>],\n            exclude_patterns=[<span class=\"hljs-string\">\"*/deprecated/*\"</span>],\n            max_links=<span class=\"hljs-number\">100</span>,\n            concurrency=<span class=\"hljs-number\">15</span>,\n            verbose=<span class=\"hljs-literal\">False</span>  <span class=\"hljs-comment\"># Clean output</span>\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://docs.example-api.com/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            api_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n\n            <span class=\"hljs-comment\"># Group by endpoint type</span>\n            endpoints = {}\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> api_links:\n                <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"head_data\"</span>):\n                    title = link[<span class=\"hljs-string\">\"head_data\"</span>].get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"\"</span>)\n                    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"GET\"</span> <span class=\"hljs-keyword\">in</span> title:\n                        endpoints.setdefault(<span class=\"hljs-string\">\"GET\"</span>, []).append(link)\n                    <span class=\"hljs-keyword\">elif</span> <span class=\"hljs-string\">\"POST\"</span> <span class=\"hljs-keyword\">in</span> title:\n                        endpoints.setdefault(<span class=\"hljs-string\">\"POST\"</span>, []).append(link)\n\n            <span class=\"hljs-keyword\">for</span> method, links <span class=\"hljs-keyword\">in</span> endpoints.items():\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{method}</span> Endpoints (<span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(links)}</span>):\"</span>)\n                <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links[:<span class=\"hljs-number\">3</span>]:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  • <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">quality_analysis</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            max_links=<span class=\"hljs-number\">200</span>,\n            concurrency=<span class=\"hljs-number\">20</span>,\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://your-website.com/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n\n            <span class=\"hljs-comment\"># Analyze intrinsic scores</span>\n            scores = [link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>) <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links]\n            avg_score = <span class=\"hljs-built_in\">sum</span>(scores) / <span class=\"hljs-built_in\">len</span>(scores) <span class=\"hljs-keyword\">if</span> scores <span class=\"hljs-keyword\">else</span> <span class=\"hljs-number\">0</span>\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📊 Link Quality Analysis:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Average intrinsic score: <span class=\"hljs-subst\">{avg_score:<span class=\"hljs-number\">.2</span>f}</span>/10.0\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   High quality links (&gt;7.0): <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>([s <span class=\"hljs-keyword\">for</span> s <span class=\"hljs-keyword\">in</span> scores <span class=\"hljs-keyword\">if</span> s &gt; <span class=\"hljs-number\">7.0</span>])}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Low quality links (&lt;3.0): <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>([s <span class=\"hljs-keyword\">for</span> s <span class=\"hljs-keyword\">in</span> scores <span class=\"hljs-keyword\">if</span> s &lt; <span class=\"hljs-number\">3.0</span>])}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Find problematic links</span>\n            bad_links = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links \n                        <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>) &lt; <span class=\"hljs-number\">2.0</span>]\n\n            <span class=\"hljs-keyword\">if</span> bad_links:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n⚠️  Links needing attention:\"</span>)\n                <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> bad_links[:<span class=\"hljs-number\">5</span>]:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span> (score: <span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>):<span class=\"hljs-number\">.1</span>f}</span>)\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"26-performance-tips\">2.6 Performance Tips</h3>\n<ol>\n<li><strong>Start Small</strong>: Begin with <code>max_links: 10</code> to understand the feature</li>\n<li><strong>Use Patterns</strong>: Filter with <code>include_patterns</code> to focus on relevant sections</li>\n<li><strong>Adjust Concurrency</strong>: Higher concurrency = faster but more resource usage</li>\n<li><strong>Set Timeouts</strong>: Use <code>timeout: 5</code> to prevent hanging on slow sites</li>\n<li><strong>Use Score Thresholds</strong>: Filter out low-quality links with <code>score_threshold</code></li>\n</ol>\n<h3 id=\"27-troubleshooting\">2.7 Troubleshooting</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Check your configuration:</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    link_preview_config<span class=\"hljs-punctuation\">=</span>LinkPreviewConfig<span class=\"hljs-punctuation\">(</span>\n        verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>   <span class=\"hljs-comment\"># ← Enable to see what's happening</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Make sure scoring is enabled:</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    score_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># ← Enable intrinsic scoring</span>\n    link_preview_config<span class=\"hljs-punctuation\">=</span>LinkPreviewConfig<span class=\"hljs-punctuation\">(</span>\n        <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"your search terms\"</span>  <span class=\"hljs-comment\"># ← For contextual scoring</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Optimize performance:</span>\nlink_preview_config = LinkPreviewConfig(\n    max_links=20,      <span class=\"hljs-comment\"># ← Reduce number</span>\n    concurrency=10,    <span class=\"hljs-comment\"># ← Increase parallelism</span>\n    <span class=\"hljs-built_in\">timeout</span>=3,         <span class=\"hljs-comment\"># ← Shorter timeout</span>\n    include_patterns=[<span class=\"hljs-string\">\"*/important/*\"</span>]  <span class=\"hljs-comment\"># ← Focus on key areas</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"3-domain-filtering\">3. Domain Filtering</h2>\n<p>Some websites contain hundreds of third-party or affiliate links. You can filter out certain domains at <strong>crawl time</strong> by configuring the crawler. The most relevant parameters in <code>CrawlerRunConfig</code> are:\n- <strong><code>exclude_external_links</code></strong>: If <code>True</code>, discard any link pointing outside the root domain.<br>\n- <strong><code>exclude_social_media_domains</code></strong>: Provide a list of social media platforms (e.g., <code>[\"facebook.com\", \"twitter.com\"]</code>) to exclude from your crawl.<br>\n- <strong><code>exclude_social_media_links</code></strong>: If <code>True</code>, automatically skip known social platforms.<br>\n- <strong><code>exclude_domains</code></strong>: Provide a list of custom domains you want to exclude (e.g., <code>[\"spammyads.com\", \"tracker.net\"]</code>).</p>\n<h3 id=\"31-example-excluding-external-social-media-links\">3.1 Example: Excluding External &amp; Social Media Links</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    crawler_cfg = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,          <span class=\"hljs-comment\"># No links outside primary domain</span>\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>       <span class=\"hljs-comment\"># Skip recognized social media domains</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            <span class=\"hljs-string\">\"https://www.example.com\"</span>,\n            config=crawler_cfg\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Crawled:\"</span>, result.url)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal links count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"External links count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])))  \n            <span class=\"hljs-comment\"># Likely zero external links in this scenario</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR]\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-example-excluding-specific-domains\">3.2 Example: Excluding Specific Domains</h3>\n<p>If you want to let external links in, but specifically exclude a domain (e.g., <code>suspiciousads.com</code>), do this:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">crawler_cfg = CrawlerRunConfig(\n    exclude_domains=[<span class=\"hljs-string\">\"suspiciousads.com\"</span>]\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"4-media-extraction\">4. Media Extraction</h2>\n<h3 id=\"41-accessing-resultmedia\">4.1 Accessing <code>result.media</code></h3>\n<p>By default, Crawl4AI collects images, audio and video URLs it finds on the page. These are stored in <code>result.media</code>, a dictionary keyed by media type (e.g., <code>images</code>, <code>videos</code>, <code>audio</code>).\n<strong>Note: Tables have been moved from <code>result.media[\"tables\"]</code> to the new <code>result.tables</code> format for better organization and direct access.</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.success:\n    <span class=\"hljs-comment\"># Get images</span>\n    images_info = result.media.get(<span class=\"hljs-string\">\"images\"</span>, [])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(images_info)}</span> images in total.\"</span>)\n    <span class=\"hljs-keyword\">for</span> i, img <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(images_info[:<span class=\"hljs-number\">3</span>]):  <span class=\"hljs-comment\"># Inspect just the first 3</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[Image <span class=\"hljs-subst\">{i}</span>] URL: <span class=\"hljs-subst\">{img[<span class=\"hljs-string\">'src'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Alt text: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'alt'</span>, <span class=\"hljs-string\">''</span>)}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Score: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'score'</span>)}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Description: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'desc'</span>, <span class=\"hljs-string\">''</span>)}</span>\\n\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">result.media = {\n  <span class=\"hljs-string\">\"images\"</span>: [\n    {\n      <span class=\"hljs-string\">\"src\"</span>: <span class=\"hljs-string\">\"https://cdn.prod.website-files.com/.../Group%2089.svg\"</span>,\n      <span class=\"hljs-string\">\"alt\"</span>: <span class=\"hljs-string\">\"coding school for kids\"</span>,\n      <span class=\"hljs-string\">\"desc\"</span>: <span class=\"hljs-string\">\"Trial Class Degrees degrees All Degrees AI Degree Technology ...\"</span>,\n      <span class=\"hljs-string\">\"score\"</span>: <span class=\"hljs-number\">3</span>,\n      <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"image\"</span>,\n      <span class=\"hljs-string\">\"group_id\"</span>: <span class=\"hljs-number\">0</span>,\n      <span class=\"hljs-string\">\"format\"</span>: <span class=\"hljs-literal\">None</span>,\n      <span class=\"hljs-string\">\"width\"</span>: <span class=\"hljs-literal\">None</span>,\n      <span class=\"hljs-string\">\"height\"</span>: <span class=\"hljs-literal\">None</span>\n    },\n    <span class=\"hljs-comment\"># ...</span>\n  ],\n  <span class=\"hljs-string\">\"videos\"</span>: [\n    <span class=\"hljs-comment\"># Similar structure but with video-specific fields</span>\n  ],\n  <span class=\"hljs-string\">\"audio\"</span>: [\n    <span class=\"hljs-comment\"># Similar structure but with audio-specific fields</span>\n  ],\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>src</code></strong>: The media URL (e.g., image source)<br>\n- <strong><code>alt</code></strong>: The alt text for images (if present)<br>\n- <strong><code>desc</code></strong>: A snippet of nearby text or a short description (optional)<br>\n- <strong><code>score</code></strong>: A heuristic relevance score if you’re using content-scoring features<br>\n- <strong><code>width</code></strong>, <strong><code>height</code></strong>: If the crawler detects dimensions for the image/video<br>\n- <strong><code>type</code></strong>: Usually <code>\"image\"</code>, <code>\"video\"</code>, or <code>\"audio\"</code><br>\n- <strong><code>group_id</code></strong>: If you’re grouping related media items, the crawler might assign an ID  <p></p>\n<h3 id=\"42-excluding-external-images\">4.2 Excluding External Images</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"43-additional-media-config\">4.3 Additional Media Config</h3>\n<ul>\n<li><strong><code>screenshot</code></strong>: Set to <code>True</code> if you want a full-page screenshot stored as <code>base64</code> in <code>result.screenshot</code>.  </li>\n<li><strong><code>pdf</code></strong>: Set to <code>True</code> if you want a PDF version of the page in <code>result.pdf</code>.  </li>\n<li><strong><code>capture_mhtml</code></strong>: Set to <code>True</code> if you want an MHTML snapshot of the page in <code>result.mhtml</code>. This format preserves the entire web page with all its resources (CSS, images, scripts) in a single file, making it perfect for archiving or offline viewing.</li>\n<li><strong><code>wait_for_images</code></strong>: If <code>True</code>, attempts to wait until images are fully loaded before final extraction.\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    crawler_cfg = CrawlerRunConfig(\n        capture_mhtml=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable MHTML capture</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=crawler_cfg)\n\n        <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.mhtml:\n            <span class=\"hljs-comment\"># Save the MHTML snapshot to a file</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"example.mhtml\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                f.write(result.mhtml)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"MHTML snapshot saved to example.mhtml\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Failed to capture MHTML:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n<li>It captures the complete page state including all resources</li>\n<li>It can be opened in most modern browsers for offline viewing</li>\n<li>It preserves the page exactly as it appeared during crawling</li>\n<li>It's a single file, making it easy to store and transfer</li>\n</ul>\n<h2 id=\"5-putting-it-all-together-link-media-filtering\">5. Putting It All Together: Link &amp; Media Filtering</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Suppose we want to keep only internal links, remove certain domains, </span>\n    <span class=\"hljs-comment\"># and discard external images from the final crawl data.</span>\n    crawler_cfg = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"spammyads.com\"</span>],\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># skip Twitter, Facebook, etc.</span>\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,      <span class=\"hljs-comment\"># keep only images from main domain</span>\n        wait_for_images=<span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># ensure images are loaded</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://www.example.com\"</span>, config=crawler_cfg)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Crawled:\"</span>, result.url)\n\n            <span class=\"hljs-comment\"># 1. Links</span>\n            in_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n            ext_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal link count:\"</span>, <span class=\"hljs-built_in\">len</span>(in_links))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"External link count:\"</span>, <span class=\"hljs-built_in\">len</span>(ext_links))  <span class=\"hljs-comment\"># should be zero with exclude_external_links=True</span>\n\n            <span class=\"hljs-comment\"># 2. Images</span>\n            images = result.media.get(<span class=\"hljs-string\">\"images\"</span>, [])\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Images found:\"</span>, <span class=\"hljs-built_in\">len</span>(images))\n\n            <span class=\"hljs-comment\"># Let's see a snippet of these images</span>\n            <span class=\"hljs-keyword\">for</span> i, img <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(images[:<span class=\"hljs-number\">3</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  - <span class=\"hljs-subst\">{img[<span class=\"hljs-string\">'src'</span>]}</span> (alt=<span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'alt'</span>,<span class=\"hljs-string\">''</span>)}</span>, score=<span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'score'</span>,<span class=\"hljs-string\">'N/A'</span>)}</span>)\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR] Failed to crawl. Reason:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"6-common-pitfalls-tips\">6. Common Pitfalls &amp; Tips</h2>\n<p>1. <strong>Conflicting Flags</strong>:<br>\n   - <code>exclude_external_links=True</code> but then also specifying <code>exclude_social_media_links=True</code> is typically fine, but understand that the first setting already discards <em>all</em> external links. The second becomes somewhat redundant.<br>\n   - <code>exclude_external_images=True</code> but want to keep some external images? Currently no partial domain-based setting for images, so you might need a custom approach or hook logic.\n2. <strong>Relevancy Scores</strong>:<br>\n   - If your version of Crawl4AI or your scraping strategy includes an <code>img[\"score\"]</code>, it’s typically a heuristic based on size, position, or content analysis. Evaluate carefully if you rely on it.\n3. <strong>Performance</strong>:<br>\n4. <strong>Social Media Lists</strong>:<br>\n   - <code>exclude_social_media_links=True</code> typically references an internal list of known social domains like Facebook, Twitter, LinkedIn, etc. If you need to add or remove from that list, look for library settings or a local config file (depending on your version).</p>\n<h1 id=\"extraction-strategies\">Extraction Strategies</h1>\n<h1 id=\"extracting-json-no-llm\">Extracting JSON (No LLM)</h1>\n<ol>\n<li><strong>Schema-based extraction</strong> with CSS or XPath selectors via <code>JsonCssExtractionStrategy</code> and <code>JsonXPathExtractionStrategy</code></li>\n<li><strong>Regular expression extraction</strong> with <code>RegexExtractionStrategy</code> for fast pattern matching</li>\n<li><strong>Faster &amp; Cheaper</strong>: No API calls or GPU overhead.  </li>\n</ol>\n<h2 id=\"1-intro-to-schema-based-extraction\">1. Intro to Schema-Based Extraction</h2>\n<ol>\n<li><strong>Nested</strong> or <strong>list</strong> types for repeated or hierarchical structures.  </li>\n</ol>\n<h2 id=\"2-simple-example-crypto-prices\">2. Simple Example: Crypto Prices</h2>\n<p>Let's begin with a <strong>simple</strong> schema-based extraction using the <code>JsonCssExtractionStrategy</code>. Below is a snippet that extracts cryptocurrency prices from a site (similar to the legacy Coinbase example). Notice we <strong>don't</strong> call any LLM:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_crypto_prices</span>():\n    <span class=\"hljs-comment\"># 1. Define a simple extraction schema</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Crypto Prices\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.crypto-row\"</span>,    <span class=\"hljs-comment\"># Repeated elements</span>\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"coin_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2.coin-name\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"price\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.coin-price\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            }\n        ]\n    }\n\n    <span class=\"hljs-comment\"># 2. Create the extraction strategy</span>\n    extraction_strategy = JsonCssExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-comment\"># 3. Set up your crawler config (if needed)</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># e.g., pass js_code or wait_for if the page is dynamic</span>\n        <span class=\"hljs-comment\"># wait_for=\"css:.crypto-row:nth-child(20)\"</span>\n        cache_mode = CacheMode.BYPASS,\n        extraction_strategy=extraction_strategy,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4. Run the crawl and extraction</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/crypto-prices\"</span>,\n\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># 5. Parse the extracted JSON</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span> coin entries\"</span>)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(data[<span class=\"hljs-number\">0</span>], indent=<span class=\"hljs-number\">2</span>) <span class=\"hljs-keyword\">if</span> data <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">\"No data found\"</span>)\n\nasyncio.run(extract_crypto_prices())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>baseSelector</code></strong>: Tells us where each \"item\" (crypto row) is.<br>\n- <strong><code>fields</code></strong>: Two fields (<code>coin_name</code>, <code>price</code>) using simple CSS selectors.<br>\n- Each field defines a <strong><code>type</code></strong> (e.g., <code>text</code>, <code>attribute</code>, <code>html</code>, <code>regex</code>, etc.).<p></p>\n<h3 id=\"xpath-example-with-raw-html\"><strong>XPath Example with <code>raw://</code> HTML</strong></h3>\n<p>Below is a short example demonstrating <strong>XPath</strong> extraction plus the <strong><code>raw://</code></strong> scheme. We'll pass a <strong>dummy HTML</strong> directly (no network request) and define the extraction strategy in <code>CrawlerRunConfig</code>.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonXPathExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_crypto_prices_xpath</span>():\n    <span class=\"hljs-comment\"># 1. Minimal dummy HTML with some repeating rows</span>\n    dummy_html = <span class=\"hljs-string\">\"\"\"\n    &lt;html&gt;\n      &lt;body&gt;\n        &lt;div class='crypto-row'&gt;\n          &lt;h2 class='coin-name'&gt;Bitcoin&lt;/h2&gt;\n          &lt;span class='coin-price'&gt;$28,000&lt;/span&gt;\n        &lt;/div&gt;\n        &lt;div class='crypto-row'&gt;\n          &lt;h2 class='coin-name'&gt;Ethereum&lt;/h2&gt;\n          &lt;span class='coin-price'&gt;$1,800&lt;/span&gt;\n        &lt;/div&gt;\n      &lt;/body&gt;\n    &lt;/html&gt;\n    \"\"\"</span>\n\n    <span class=\"hljs-comment\"># 2. Define the JSON schema (XPath version)</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Crypto Prices via XPath\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"//div[@class='crypto-row']\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"coin_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".//h2[@class='coin-name']\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"price\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".//span[@class='coin-price']\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            }\n        ]\n    }\n\n    <span class=\"hljs-comment\"># 3. Place the strategy in the CrawlerRunConfig</span>\n    config = CrawlerRunConfig(\n        extraction_strategy=JsonXPathExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n    )\n\n    <span class=\"hljs-comment\"># 4. Use raw:// scheme to pass dummy_html directly</span>\n    raw_url = <span class=\"hljs-string\">f\"raw://<span class=\"hljs-subst\">{dummy_html}</span>\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=raw_url,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span> coin rows\"</span>)\n        <span class=\"hljs-keyword\">if</span> data:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"First item:\"</span>, data[<span class=\"hljs-number\">0</span>])\n\nasyncio.run(extract_crypto_prices_xpath())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. <strong><code>JsonXPathExtractionStrategy</code></strong> is used instead of <code>JsonCssExtractionStrategy</code>.<br>\n2. <strong><code>baseSelector</code></strong> and each field's <code>\"selector\"</code> use <strong>XPath</strong> instead of CSS.<br>\n3. <strong><code>raw://</code></strong> lets us pass <code>dummy_html</code> with no real network request—handy for local testing.<br>\n4. Everything (including the extraction strategy) is in <strong><code>CrawlerRunConfig</code></strong>.<br>\nThat's how you keep the config self-contained, illustrate <strong>XPath</strong> usage, and demonstrate the <strong>raw</strong> scheme for direct HTML input—all while avoiding the old approach of passing <code>extraction_strategy</code> directly to <code>arun()</code>.<p></p>\n<h2 id=\"3-advanced-schema-nested-structures\">3. Advanced Schema &amp; Nested Structures</h2>\n<h3 id=\"sample-e-commerce-html\">Sample E-Commerce HTML</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/examples/sample_ecommerce.html\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"E-commerce Product Catalog\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.category\"</span>,\n    <span class=\"hljs-comment\"># (1) We can define optional baseFields if we want to extract attributes </span>\n    <span class=\"hljs-comment\"># from the category container</span>\n    <span class=\"hljs-string\">\"baseFields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"data_cat_id\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"data-cat-id\"</span><span class=\"hljs-punctuation\">}</span>, \n    <span class=\"hljs-punctuation\">]</span>,\n    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span>\n            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"category_name\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h2.category-name\"</span>,\n            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n        <span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span>\n            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"products\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.product\"</span>,\n            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested_list\"</span>,    <span class=\"hljs-comment\"># repeated sub-objects</span>\n            <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"name\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h3.product-name\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"price\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.product-price\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"details\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.product-details\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested\"</span>,  <span class=\"hljs-comment\"># single sub-object</span>\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"brand\"</span>,\n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.brand\"</span>,\n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"model\"</span>,\n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.model\"</span>,\n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"features\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"ul.product-features li\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"feature\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span> \n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"reviews\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.review\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested_list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"reviewer\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.reviewer\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"rating\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.rating\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"comment\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.review-text\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"related_products\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"ul.related-products li\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"name\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.related-name\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"price\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.related-price\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>\n            <span class=\"hljs-punctuation\">]</span>\n        <span class=\"hljs-punctuation\">}</span>\n    <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong>Nested vs. List</strong>:<br>\n  - <strong><code>type: \"nested\"</code></strong> means a <strong>single</strong> sub-object (like <code>details</code>).<br>\n  - <strong><code>type: \"list\"</code></strong> means multiple items that are <strong>simple</strong> dictionaries or single text fields.<br>\n  - <strong><code>type: \"nested_list\"</code></strong> means repeated <strong>complex</strong> objects (like <code>products</code> or <code>reviews</code>).\n- <strong>Base Fields</strong>: We can extract <strong>attributes</strong> from the container element via <code>\"baseFields\"</code>. For instance, <code>\"data_cat_id\"</code> might be <code>data-cat-id=\"elect123\"</code>.<br>\n- <strong>Transforms</strong>: We can also define a <code>transform</code> if we want to lower/upper case, strip whitespace, or even run a custom function.<p></p>\n<h3 id=\"running-the-extraction\">Running the Extraction</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\necommerce_schema = {\n    <span class=\"hljs-comment\"># ... the advanced schema from above ...</span>\n}\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_ecommerce_data</span>():\n    strategy = JsonCssExtractionStrategy(ecommerce_schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n    config = CrawlerRunConfig()\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/examples/sample_ecommerce.html\"</span>,\n            extraction_strategy=strategy,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># Parse the JSON output</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(data, indent=<span class=\"hljs-number\">2</span>) <span class=\"hljs-keyword\">if</span> data <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">\"No data found.\"</span>)\n\nasyncio.run(extract_ecommerce_data())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nIf all goes well, you get a <strong>structured</strong> JSON array with each \"category,\" containing an array of <code>products</code>. Each product includes <code>details</code>, <code>features</code>, <code>reviews</code>, etc. All of that <strong>without</strong> an LLM.<p></p>\n<h2 id=\"4-regexextractionstrategy-fast-pattern-based-extraction\">4. RegexExtractionStrategy - Fast Pattern-Based Extraction</h2>\n<p>Crawl4AI now offers a powerful new zero-LLM extraction strategy: <code>RegexExtractionStrategy</code>. This strategy provides lightning-fast extraction of common data types like emails, phone numbers, URLs, dates, and more using pre-compiled regular expressions.</p>\n<h3 id=\"key-features\">Key Features</h3>\n<ul>\n<li><strong>Zero LLM Dependency</strong>: Extracts data without any AI model calls</li>\n<li><strong>Blazing Fast</strong>: Uses pre-compiled regex patterns for maximum performance</li>\n<li><strong>Built-in Patterns</strong>: Includes ready-to-use patterns for common data types</li>\n</ul>\n<h3 id=\"simple-example-extracting-common-entities\">Simple Example: Extracting Common Entities</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_with_regex</span>():\n    <span class=\"hljs-comment\"># Create a strategy using built-in patterns for URLs and currencies</span>\n    strategy = RegexExtractionStrategy(\n        pattern = RegexExtractionStrategy.Url | RegexExtractionStrategy.Currency\n    )\n\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data[:<span class=\"hljs-number\">5</span>]:  <span class=\"hljs-comment\"># Show first 5 matches</span>\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'label'</span>]}</span>: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total matches: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span>\"</span>)\n\nasyncio.run(extract_with_regex())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"available-built-in-patterns\">Available Built-in Patterns</h3>\n<p><code>RegexExtractionStrategy</code> provides these common patterns as IntFlag attributes for easy combining:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Use individual patterns</span>\nstrategy = RegexExtractionStrategy(pattern=RegexExtractionStrategy.Email)\n\n<span class=\"hljs-comment\"># Combine multiple patterns</span>\nstrategy = RegexExtractionStrategy(\n    pattern = (\n        RegexExtractionStrategy.Email | \n        RegexExtractionStrategy.PhoneUS | \n        RegexExtractionStrategy.Url\n    )\n)\n\n<span class=\"hljs-comment\"># Use all available patterns</span>\nstrategy = RegexExtractionStrategy(pattern=RegexExtractionStrategy.All)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>Email</code> - Email addresses\n- <code>PhoneIntl</code> - International phone numbers\n- <code>PhoneUS</code> - US-format phone numbers\n- <code>Url</code> - HTTP/HTTPS URLs\n- <code>IPv4</code> - IPv4 addresses\n- <code>IPv6</code> - IPv6 addresses\n- <code>Uuid</code> - UUIDs\n- <code>Currency</code> - Currency values (USD, EUR, etc.)\n- <code>Percentage</code> - Percentage values\n- <code>Number</code> - Numeric values\n- <code>DateIso</code> - ISO format dates\n- <code>DateUS</code> - US format dates\n- <code>Time24h</code> - 24-hour format times\n- <code>PostalUS</code> - US postal codes\n- <code>PostalUK</code> - UK postal codes\n- <code>HexColor</code> - HTML hex color codes\n- <code>TwitterHandle</code> - Twitter handles\n- <code>Hashtag</code> - Hashtags\n- <code>MacAddr</code> - MAC addresses\n- <code>Iban</code> - International bank account numbers\n- <code>CreditCard</code> - Credit card numbers<p></p>\n<h3 id=\"custom-pattern-example\">Custom Pattern Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_prices</span>():\n    <span class=\"hljs-comment\"># Define a custom pattern for US Dollar prices</span>\n    price_pattern = {<span class=\"hljs-string\">\"usd_price\"</span>: <span class=\"hljs-string\">r\"\\$\\s?\\d{1,3}(?:,\\d{3})*(?:\\.\\d{2})?\"</span>}\n\n    <span class=\"hljs-comment\"># Create strategy with custom pattern</span>\n    strategy = RegexExtractionStrategy(custom=price_pattern)\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.example.com/products\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found price: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n\nasyncio.run(extract_prices())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"llm-assisted-pattern-generation\">LLM-Assisted Pattern Generation</h3>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> pathlib <span class=\"hljs-keyword\">import</span> Path\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy,\n    LLMConfig\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_with_generated_pattern</span>():\n    cache_dir = Path(<span class=\"hljs-string\">\"./pattern_cache\"</span>)\n    cache_dir.mkdir(exist_ok=<span class=\"hljs-literal\">True</span>)\n    pattern_file = cache_dir / <span class=\"hljs-string\">\"price_pattern.json\"</span>\n\n    <span class=\"hljs-comment\"># 1. Generate or load pattern</span>\n    <span class=\"hljs-keyword\">if</span> pattern_file.exists():\n        pattern = json.load(pattern_file.<span class=\"hljs-built_in\">open</span>())\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Using cached pattern: <span class=\"hljs-subst\">{pattern}</span>\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Generating pattern via LLM...\"</span>)\n\n        <span class=\"hljs-comment\"># Configure LLM</span>\n        llm_config = LLMConfig(\n            provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>,\n            api_token=<span class=\"hljs-string\">\"env:OPENAI_API_KEY\"</span>,\n        )\n\n        <span class=\"hljs-comment\"># Get sample HTML for context</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/products\"</span>)\n            html = result.markdown.fit_html\n\n        <span class=\"hljs-comment\"># Generate pattern (one-time LLM usage)</span>\n        pattern = RegexExtractionStrategy.generate_pattern(\n            label=<span class=\"hljs-string\">\"price\"</span>,\n            html=html,\n            query=<span class=\"hljs-string\">\"Product prices in USD format\"</span>,\n            llm_config=llm_config,\n        )\n\n        <span class=\"hljs-comment\"># Cache pattern for future use</span>\n        json.dump(pattern, pattern_file.<span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"w\"</span>), indent=<span class=\"hljs-number\">2</span>)\n\n    <span class=\"hljs-comment\"># 2. Use pattern for extraction (no LLM calls)</span>\n    strategy = RegexExtractionStrategy(custom=pattern)\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data[:<span class=\"hljs-number\">10</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total matches: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span>\"</span>)\n\nasyncio.run(extract_with_generated_pattern())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. Use an LLM once to generate a highly optimized regex for your specific site\n2. Save the pattern to disk for reuse \n3. Extract data using only regex (no further LLM calls) in production<p></p>\n<h3 id=\"extraction-results-format\">Extraction Results Format</h3>\n<p>The <code>RegexExtractionStrategy</code> returns results in a consistent format:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">[\n  {\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-string\">\"label\"</span>: <span class=\"hljs-string\">\"email\"</span>,\n    <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"contact@example.com\"</span>,\n    <span class=\"hljs-string\">\"span\"</span>: [<span class=\"hljs-number\">145</span>, <span class=\"hljs-number\">163</span>]\n  },\n  {\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-string\">\"label\"</span>: <span class=\"hljs-string\">\"url\"</span>,\n    <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"https://support.example.com\"</span>,\n    <span class=\"hljs-string\">\"span\"</span>: [<span class=\"hljs-number\">210</span>, <span class=\"hljs-number\">235</span>]\n  }\n]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <code>url</code>: The source URL\n- <code>label</code>: The pattern name that matched (e.g., \"email\", \"phone_us\")\n- <code>value</code>: The extracted text\n- <code>span</code>: The start and end positions in the source content<p></p>\n<h2 id=\"5-why-no-llm-is-often-better\">5. Why \"No LLM\" Is Often Better</h2>\n<h2 id=\"6-base-element-attributes-additional-fields\">6. Base Element Attributes &amp; Additional Fields</h2>\n<p>It's easy to <strong>extract attributes</strong> (like <code>href</code>, <code>src</code>, or <code>data-xxx</code>) from your base or nested elements using:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\"><span class=\"hljs-keyword\">null</span></span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nYou can define them in <strong><code>baseFields</code></strong> (extracted from the main container element) or in each field's sub-lists. This is especially helpful if you need an item's link or ID stored in the parent <code>&lt;div&gt;</code>.<p></p>\n<h2 id=\"7-putting-it-all-together-larger-example\">7. Putting It All Together: Larger Example</h2>\n<p>Consider a blog site. We have a schema that extracts the <strong>URL</strong> from each post card (via <code>baseFields</code> with an <code>\"attribute\": \"href\"</code>), plus the title, date, summary, and author:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Blog Posts\"</span>,\n  <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"a.blog-post-card\"</span>,\n  <span class=\"hljs-string\">\"baseFields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"post_url\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">]</span>,\n  <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h2.post-title\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"No Title\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"date\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"time.post-date\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"summary\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.post-summary\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"author\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.post-author\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThen run with <code>JsonCssExtractionStrategy(schema)</code> to get an array of blog post objects, each with <code>\"post_url\"</code>, <code>\"title\"</code>, <code>\"date\"</code>, <code>\"summary\"</code>, <code>\"author\"</code>.<p></p>\n<h2 id=\"8-tips-best-practices\">8. Tips &amp; Best Practices</h2>\n<ol>\n<li><strong>Test</strong> your schema on partial HTML or a test page before a big crawl.  </li>\n<li><strong>Combine with JS Execution</strong> if the site loads content dynamically. You can pass <code>js_code</code> or <code>wait_for</code> in <code>CrawlerRunConfig</code>.  </li>\n<li><strong>Look at Logs</strong> when <code>verbose=True</code>: if your selectors are off or your schema is malformed, it'll often show warnings.  </li>\n<li><strong>Use baseFields</strong> if you need attributes from the container element (e.g., <code>href</code>, <code>data-id</code>), especially for the \"parent\" item.  </li>\n<li><strong>Consider Using Regex First</strong>: For simple data types like emails, URLs, and dates, <code>RegexExtractionStrategy</code> is often the fastest approach.</li>\n</ol>\n<h2 id=\"9-schema-generation-utility\">9. Schema Generation Utility</h2>\n<ol>\n<li>You're dealing with a new website structure and want a quick starting point</li>\n<li>You need to extract complex nested data structures</li>\n<li>You want to avoid the learning curve of CSS/XPath selector syntax</li>\n</ol>\n<h3 id=\"using-the-schema-generator\">Using the Schema Generator</h3>\n<p>The schema generator is available as a static method on both <code>JsonCssExtractionStrategy</code> and <code>JsonXPathExtractionStrategy</code>. You can choose between OpenAI's GPT-4 or the open-source Ollama for schema generation:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy, JsonXPathExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Sample HTML with product information</span>\nhtml = <span class=\"hljs-string\">\"\"\"\n&lt;div class=\"product-card\"&gt;\n    &lt;h2 class=\"title\"&gt;Gaming Laptop&lt;/h2&gt;\n    &lt;div class=\"price\"&gt;$999.99&lt;/div&gt;\n    &lt;div class=\"specs\"&gt;\n        &lt;ul&gt;\n            &lt;li&gt;16GB RAM&lt;/li&gt;\n            &lt;li&gt;1TB SSD&lt;/li&gt;\n        &lt;/ul&gt;\n    &lt;/div&gt;\n&lt;/div&gt;\n\"\"\"</span>\n\n<span class=\"hljs-comment\"># Option 1: Using OpenAI (requires API token)</span>\ncss_schema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    schema_type=<span class=\"hljs-string\">\"css\"</span>, \n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-openai-token\"</span>)\n)\n\n<span class=\"hljs-comment\"># Option 2: Using Ollama (open source, no token needed)</span>\nxpath_schema = JsonXPathExtractionStrategy.generate_schema(\n    html,\n    schema_type=<span class=\"hljs-string\">\"xpath\"</span>,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama3.3\"</span>, api_token=<span class=\"hljs-literal\">None</span>)  <span class=\"hljs-comment\"># Not needed for Ollama</span>\n)\n\n<span class=\"hljs-comment\"># Use the generated schema for fast, repeated extractions</span>\nstrategy = JsonCssExtractionStrategy(css_schema)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"token-usage-tracking\">Token Usage Tracking</h3>\n<p><code>generate_schema</code> may make multiple LLM calls internally (field inference, generation, validation retries). Track total token consumption by passing a <code>TokenUsage</code> accumulator:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.models <span class=\"hljs-keyword\">import</span> TokenUsage\n\nusage = TokenUsage()\nschema = JsonCssExtractionStrategy.generate_schema(\n    url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n    query=<span class=\"hljs-string\">\"extract product name and price\"</span>,\n    usage=usage,\n)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total tokens: <span class=\"hljs-subst\">{usage.total_tokens}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThe <code>usage</code> parameter is optional and fully backward-compatible. Both <code>generate_schema</code> (sync) and <code>agenerate_schema</code> (async) support it.<p></p>\n<h3 id=\"llm-provider-options\">LLM Provider Options</h3>\n<ol>\n<li><strong>OpenAI GPT-4 (<code>openai/gpt4o</code>)</strong></li>\n<li>Default provider</li>\n<li>Requires an API token</li>\n<li>Generally provides more accurate schemas</li>\n<li>Set via environment variable: <code>OPENAI_API_KEY</code></li>\n<li><strong>Ollama (<code>ollama/llama3.3</code>)</strong></li>\n<li>Open source alternative</li>\n<li>No API token required</li>\n<li>Self-hosted option</li>\n<li>Good for development and testing</li>\n</ol>\n<h3 id=\"benefits-of-schema-generation\">Benefits of Schema Generation</h3>\n<h3 id=\"best-practices\">Best Practices</h3>\n<ol>\n<li><strong>Choose Provider Wisely</strong>: </li>\n<li>Use OpenAI for production-quality schemas</li>\n<li>Use Ollama for development, testing, or when you need a self-hosted solution</li>\n</ol>\n<h2 id=\"10-conclusion_1\">10. Conclusion</h2>\n<p>With Crawl4AI's LLM-free extraction strategies - <code>JsonCssExtractionStrategy</code>, <code>JsonXPathExtractionStrategy</code>, and now <code>RegexExtractionStrategy</code> - you can build powerful pipelines that:\n- Scrape any consistent site for structured data.<br>\n- Support nested objects, repeating lists, or pattern-based extraction.<br>\n- Scale to thousands of pages quickly and reliably.\n- Use <strong><code>RegexExtractionStrategy</code></strong> for fast extraction of common data types like emails, phones, URLs, dates, etc.\n- Use <strong><code>JsonCssExtractionStrategy</code></strong> or <strong><code>JsonXPathExtractionStrategy</code></strong> for structured data with clear HTML patterns</p>\n<h1 id=\"extracting-json-llm\">Extracting JSON (LLM)</h1>\n<p><strong>Important</strong>: LLM-based extraction can be slower and costlier than schema-based approaches. If your page data is highly structured, consider using <a href=\"./no-llm-strategies.md\"><code>JsonCssExtractionStrategy</code></a> or <a href=\"./no-llm-strategies.md\"><code>JsonXPathExtractionStrategy</code></a> first. But if you need AI to interpret or reorganize content, read on!</p>\n<h2 id=\"1-why-use-an-llm\">1. Why Use an LLM?</h2>\n<h2 id=\"2-provider-agnostic-via-litellm\">2. Provider-Agnostic via LiteLLM</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">llmConfig</span> = LlmConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>, api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nCrawl4AI uses a “provider string” (e.g., <code>\"openai/gpt-4o\"</code>, <code>\"ollama/llama2.0\"</code>, <code>\"aws/titan\"</code>) to identify your LLM. <strong>Any</strong> model that LiteLLM supports is fair game. You just provide:\n- <strong><code>provider</code></strong>: The <code>&lt;provider&gt;/&lt;model_name&gt;</code> identifier (e.g., <code>\"openai/gpt-4\"</code>, <code>\"ollama/llama2\"</code>, <code>\"huggingface/google-flan\"</code>, etc.).<br>\n- <strong><code>api_token</code></strong>: If needed (for OpenAI, HuggingFace, etc.); local models or Ollama might not require it.<br>\n- <strong><code>base_url</code></strong> (optional): If your provider has a custom endpoint.  <p></p>\n<h2 id=\"3-how-llm-extraction-works\">3. How LLM Extraction Works</h2>\n<h3 id=\"31-flow\">3.1 Flow</h3>\n<p>1. <strong>Chunking</strong> (optional): The HTML or markdown is split into smaller segments if it’s very long (based on <code>chunk_token_threshold</code>, overlap, etc.).<br>\n2. <strong>Prompt Construction</strong>: For each chunk, the library forms a prompt that includes your <strong><code>instruction</code></strong> (and possibly schema or examples).<br>\n4. <strong>Combining</strong>: The results from each chunk are merged and parsed into JSON.</p>\n<h3 id=\"32-extraction_type\">3.2 <code>extraction_type</code></h3>\n<ul>\n<li><strong><code>\"schema\"</code></strong>: The model tries to return JSON conforming to your Pydantic-based schema.  </li>\n<li><strong><code>\"block\"</code></strong>: The model returns freeform text, or smaller JSON structures, which the library collects.<br>\nFor structured data, <code>\"schema\"</code> is recommended. You provide <code>schema=YourPydanticModel.model_json_schema()</code>.</li>\n</ul>\n<h2 id=\"4-key-parameters\">4. Key Parameters</h2>\n<p>Below is an overview of important LLM extraction parameters. All are typically set inside <code>LLMExtractionStrategy(...)</code>. You then put that strategy in your <code>CrawlerRunConfig(..., extraction_strategy=...)</code>.\n1. <strong><code>llmConfig</code></strong> (LlmConfig): e.g., <code>\"openai/gpt-4\"</code>, <code>\"ollama/llama2\"</code>.  <br>\n2. <strong><code>schema</code></strong> (dict): A JSON schema describing the fields you want. Usually generated by <code>YourModel.model_json_schema()</code>.<br>\n3. <strong><code>extraction_type</code></strong> (str): <code>\"schema\"</code> or <code>\"block\"</code>.<br>\n4. <strong><code>instruction</code></strong> (str): Prompt text telling the LLM what you want extracted. E.g., “Extract these fields as a JSON array.”<br>\n5. <strong><code>chunk_token_threshold</code></strong> (int): Maximum tokens per chunk. If your content is huge, you can break it up for the LLM.<br>\n6. <strong><code>overlap_rate</code></strong> (float): Overlap ratio between adjacent chunks. E.g., <code>0.1</code> means 10% of each chunk is repeated to preserve context continuity.<br>\n7. <strong><code>apply_chunking</code></strong> (bool): Set <code>True</code> to chunk automatically. If you want a single pass, set <code>False</code>.<br>\n8. <strong><code>input_format</code></strong> (str): Determines <strong>which</strong> crawler result is passed to the LLM. Options include:<br>\n   - <code>\"markdown\"</code>: The raw markdown (default).<br>\n   - <code>\"fit_markdown\"</code>: The filtered “fit” markdown if you used a content filter.<br>\n   - <code>\"html\"</code>: The cleaned or raw HTML.<br>\n9. <strong><code>extra_args</code></strong> (dict): Additional LLM parameters like <code>temperature</code>, <code>max_tokens</code>, <code>top_p</code>, etc.<br>\n10. <strong><code>show_usage()</code></strong>: A method you can call to print out usage info (token usage per chunk, total cost if known).<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">extraction_strategy <span class=\"hljs-punctuation\">=</span> LLMExtractionStrategy<span class=\"hljs-punctuation\">(</span>\n    llm_config <span class=\"hljs-punctuation\">=</span> LLMConfig<span class=\"hljs-punctuation\">(</span>provider<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"openai/gpt-4\"</span>, api_token<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"YOUR_OPENAI_KEY\"</span><span class=\"hljs-punctuation\">)</span>,\n    <span class=\"hljs-keyword\">schema</span><span class=\"hljs-punctuation\">=</span>MyModel.model_json_schema<span class=\"hljs-punctuation\">(</span><span class=\"hljs-punctuation\">)</span>,\n    extraction_type<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"schema\"</span>,\n    instruction<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Extract a list of items from the text with 'name' and 'price' fields.\"</span>,\n    chunk_token_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">1200</span>,\n    overlap_rate<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.1</span>,\n    apply_chunking<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    input_format<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"html\"</span>,\n    extra_args<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"temperature\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">0.1</span>, <span class=\"hljs-string\">\"max_tokens\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1000</span><span class=\"hljs-punctuation\">}</span>,\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"5-putting-it-in-crawlerrunconfig\">5. Putting It in <code>CrawlerRunConfig</code></h2>\n<p><strong>Important</strong>: In Crawl4AI, all strategy definitions should go inside the <code>CrawlerRunConfig</code>, not directly as a param in <code>arun()</code>. Here’s a full example:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> typing <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">List</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Product</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    name: <span class=\"hljs-built_in\">str</span>\n    price: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Define the LLM extraction strategy</span>\n    llm_strategy = LLMExtractionStrategy(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>, api_token=os.getenv(<span class=\"hljs-string\">'OPENAI_API_KEY'</span>)),\n        schema=Product.schema_json(), <span class=\"hljs-comment\"># Or use model_json_schema()</span>\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract all product objects with 'name' and 'price' from the content.\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">1000</span>,\n        overlap_rate=<span class=\"hljs-number\">0.0</span>,\n        apply_chunking=<span class=\"hljs-literal\">True</span>,\n        input_format=<span class=\"hljs-string\">\"markdown\"</span>,   <span class=\"hljs-comment\"># or \"html\", \"fit_markdown\"</span>\n        extra_args={<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0.0</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">800</span>}\n    )\n\n    <span class=\"hljs-comment\"># 2. Build the crawler config</span>\n    crawl_config = CrawlerRunConfig(\n        extraction_strategy=llm_strategy,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-comment\"># 3. Create a browser config if needed</span>\n    browser_cfg = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4. Let's say we want to crawl a single page</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n            config=crawl_config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># 5. The extracted content is presumably JSON</span>\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted items:\"</span>, data)\n\n            <span class=\"hljs-comment\"># 6. Show usage stats</span>\n            llm_strategy.show_usage()  <span class=\"hljs-comment\"># prints token usage</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"6-chunking-details\">6. Chunking Details</h2>\n<h3 id=\"61-chunk_token_threshold\">6.1 <code>chunk_token_threshold</code></h3>\n<p>If your page is large, you might exceed your LLM’s context window. <strong><code>chunk_token_threshold</code></strong> sets the approximate max tokens per chunk. The library calculates word→token ratio using <code>word_token_rate</code> (often ~0.75 by default). If chunking is enabled (<code>apply_chunking=True</code>), the text is split into segments.</p>\n<h3 id=\"62-overlap_rate\">6.2 <code>overlap_rate</code></h3>\n<p>To keep context continuous across chunks, we can overlap them. E.g., <code>overlap_rate=0.1</code> means each subsequent chunk includes 10% of the previous chunk’s text. This is helpful if your needed info might straddle chunk boundaries.</p>\n<h3 id=\"63-performance-parallelism\">6.3 Performance &amp; Parallelism</h3>\n<h2 id=\"7-input-format\">7. Input Format</h2>\n<p>By default, <strong>LLMExtractionStrategy</strong> uses <code>input_format=\"markdown\"</code>, meaning the <strong>crawler’s final markdown</strong> is fed to the LLM. You can change to:\n- <strong><code>html</code></strong>: The cleaned HTML or raw HTML (depending on your crawler config) goes into the LLM.<br>\n- <strong><code>fit_markdown</code></strong>: If you used, for instance, <code>PruningContentFilter</code>, the “fit” version of the markdown is used. This can drastically reduce tokens if you trust the filter.<br>\n- <strong><code>markdown</code></strong>: Standard markdown output from the crawler’s <code>markdown_generator</code>.\nThis setting is crucial: if the LLM instructions rely on HTML tags, pick <code>\"html\"</code>. If you prefer a text-based approach, pick <code>\"markdown\"</code>.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">LLMExtractionStrategy(\n    <span class=\"hljs-comment\"># ...</span>\n    input_format=<span class=\"hljs-string\">\"html\"</span>,  <span class=\"hljs-comment\"># Instead of \"markdown\" or \"fit_markdown\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"8-token-usage-show-usage\">8. Token Usage &amp; Show Usage</h2>\n<ul>\n<li><strong><code>usages</code></strong> (list): token usage per chunk or call.  </li>\n<li><strong><code>total_usage</code></strong>: sum of all chunk calls.  </li>\n<li><strong><code>show_usage()</code></strong>: prints a usage report (if the provider returns usage data).\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">llm_strategy = LLMExtractionStrategy(...)\n<span class=\"hljs-comment\"># ...</span>\nllm_strategy.show_usage()\n<span class=\"hljs-comment\"># e.g. “Total usage: 1241 tokens across 2 chunk calls”</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h2 id=\"9-example-building-a-knowledge-graph\">9. Example: Building a Knowledge Graph</h2>\n<p>Below is a snippet combining <strong><code>LLMExtractionStrategy</code></strong> with a Pydantic schema for a knowledge graph. Notice how we pass an <strong><code>instruction</code></strong> telling the model what to parse.\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> typing <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">List</span>\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Entity</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    name: <span class=\"hljs-built_in\">str</span>\n    description: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Relationship</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    entity1: Entity\n    entity2: Entity\n    description: <span class=\"hljs-built_in\">str</span>\n    relation_type: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">KnowledgeGraph</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    entities: <span class=\"hljs-type\">List</span>[Entity]\n    relationships: <span class=\"hljs-type\">List</span>[Relationship]\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># LLM extraction strategy</span>\n    llm_strat = LLMExtractionStrategy(\n        llmConfig = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4\"</span>, api_token=os.getenv(<span class=\"hljs-string\">'OPENAI_API_KEY'</span>)),\n        schema=KnowledgeGraph.model_json_schema(),\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract entities and relationships from the content. Return valid JSON.\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">1400</span>,\n        apply_chunking=<span class=\"hljs-literal\">True</span>,\n        input_format=<span class=\"hljs-string\">\"html\"</span>,\n        extra_args={<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0.1</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">1500</span>}\n    )\n\n    crawl_config = CrawlerRunConfig(\n        extraction_strategy=llm_strat,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Example page</span>\n        url = <span class=\"hljs-string\">\"https://www.nbcnews.com/business\"</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=crawl_config)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"--- LLM RAW RESPONSE ---\"</span>)\n        <span class=\"hljs-built_in\">print</span>(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"--- END LLM RAW RESPONSE ---\"</span>)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"kb_result.json\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                f.write(result.extracted_content)\n            llm_strat.show_usage()\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n- <strong><code>extraction_type=\"schema\"</code></strong> ensures we get JSON fitting our <code>KnowledgeGraph</code>.<br>\n- <strong><code>input_format=\"html\"</code></strong> means we feed HTML to the model.<br>\n- <strong><code>instruction</code></strong> guides the model to output a structured knowledge graph.  <p></p>\n<h2 id=\"10-best-practices-caveats\">10. Best Practices &amp; Caveats</h2>\n<p>4. <strong>Schema Strictness</strong>: <code>\"schema\"</code> extraction tries to parse the model output as JSON. If the model returns invalid JSON, partial extraction might happen, or you might get an error.  </p>\n<h2 id=\"11-conclusion\">11. Conclusion</h2>\n<ul>\n<li>Put your LLM strategy <strong>in <code>CrawlerRunConfig</code></strong>.  </li>\n<li>Use <strong><code>input_format</code></strong> to pick which form (markdown, HTML, fit_markdown) the LLM sees.  </li>\n<li>Tweak <strong><code>chunk_token_threshold</code></strong>, <strong><code>overlap_rate</code></strong>, and <strong><code>apply_chunking</code></strong> to handle large content efficiently.  </li>\n<li>Monitor token usage with <code>show_usage()</code>.\nIf your site’s data is consistent or repetitive, consider <a href=\"./no-llm-strategies.md\"><code>JsonCssExtractionStrategy</code></a> first for speed and simplicity. But if you need an <strong>AI-driven</strong> approach, <code>LLMExtractionStrategy</code> offers a flexible, multi-provider solution for extracting structured JSON from any website.\n1. <strong>Experiment with Different Providers</strong>  </li>\n<li>Try switching the <code>provider</code> (e.g., <code>\"ollama/llama2\"</code>, <code>\"openai/gpt-4o\"</code>, etc.) to see differences in speed, accuracy, or cost.  </li>\n<li>Pass different <code>extra_args</code> like <code>temperature</code>, <code>top_p</code>, and <code>max_tokens</code> to fine-tune your results.\n2. <strong>Performance Tuning</strong>  </li>\n<li>If pages are large, tweak <code>chunk_token_threshold</code>, <code>overlap_rate</code>, or <code>apply_chunking</code> to optimize throughput.  </li>\n<li>Check the usage logs with <code>show_usage()</code> to keep an eye on token consumption and identify potential bottlenecks.\n3. <strong>Validate Outputs</strong>  </li>\n<li>If using <code>extraction_type=\"schema\"</code>, parse the LLM’s JSON with a Pydantic model for a final validation step.<br>\n4. <strong>Explore Hooks &amp; Automation</strong>  </li>\n</ul>\n<h1 id=\"advanced-features\">Advanced Features</h1>\n<h1 id=\"session-management\">Session Management</h1>\n<ul>\n<li><strong>Performing JavaScript actions before and after crawling.</strong>\nUse <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> to maintain state with a <code>session_id</code>:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs import BrowserConfig, <span class=\"hljs-function\">CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    session_id</span> = <span class=\"hljs-string\">\"my_session\"</span>\n\n    <span class=\"hljs-meta\"># Define configurations</span>\n    config1 = CrawlerRunConfig(\n        url=<span class=\"hljs-string\">\"https://example.com/page1\"</span>, session_id=session_id\n    )\n    config2 = CrawlerRunConfig(\n        url=<span class=\"hljs-string\">\"https://example.com/page2\"</span>, session_id=session_id\n    )\n\n    <span class=\"hljs-meta\"># First request</span>\n    result1 = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config1)\n\n    <span class=\"hljs-meta\"># Subsequent request using the same session</span>\n    result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config2)\n\n    <span class=\"hljs-meta\"># Clean up when done</span>\n    <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.cache_context <span class=\"hljs-keyword\">import</span> CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_dynamic_content</span>():\n    url = <span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>\n    session_id = <span class=\"hljs-string\">\"wait_for_session\"</span>\n    all_commits = []\n\n    js_next_page = <span class=\"hljs-string\">\"\"\"\n    const commits = document.querySelectorAll('li[data-testid=\"commit-row-item\"] h4');\n    if (commits.length &gt; 0) {\n        window.lastCommit = commits[0].textContent.trim();\n    }\n    const button = document.querySelector('a[data-testid=\"pagination-next-button\"]');\n    if (button) {button.click(); console.log('button clicked') }\n    \"\"\"</span>\n\n    wait_for = <span class=\"hljs-string\">\"\"\"() =&gt; {\n        const commits = document.querySelectorAll('li[data-testid=\"commit-row-item\"] h4');\n        if (commits.length === 0) return false;\n        const firstCommit = commits[0].textContent.trim();\n        return firstCommit !== window.lastCommit;\n    }\"\"\"</span>\n\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Commit Extractor\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"li[data-testid='commit-row-item']\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h4 a\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n                <span class=\"hljs-string\">\"transform\"</span>: <span class=\"hljs-string\">\"strip\"</span>,\n            },\n        ],\n    }\n    extraction_strategy = JsonCssExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n    browser_config = BrowserConfig(\n        verbose=<span class=\"hljs-literal\">True</span>,\n        headless=<span class=\"hljs-literal\">False</span>,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            crawler_config = CrawlerRunConfig(\n                session_id=session_id,\n                css_selector=<span class=\"hljs-string\">\"li[data-testid='commit-row-item']\"</span>,\n                extraction_strategy=extraction_strategy,\n                js_code=js_next_page <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                wait_for=wait_for <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS,\n                capture_console_messages=<span class=\"hljs-literal\">True</span>,\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=crawler_config)\n\n            <span class=\"hljs-keyword\">if</span> result.console_messages:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span> console messages:\"</span>, result.console_messages)\n\n            <span class=\"hljs-keyword\">if</span> result.extracted_content:\n                <span class=\"hljs-comment\"># print(f\"Page {page + 1} result:\", result.extracted_content)</span>\n                commits = json.loads(result.extracted_content)\n                all_commits.extend(commits)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(commits)}</span> commits\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: No content extracted\"</span>)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(all_commits)}</span> commits across 3 pages\"</span>)\n        <span class=\"hljs-comment\"># Clean up session</span>\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ul>\n<h2 id=\"example-1-basic-session-based-crawling\">Example 1: Basic Session-Based Crawling</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.cache_context <span class=\"hljs-keyword\">import</span> CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">basic_session_crawl</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"dynamic_content_session\"</span>\n        url = <span class=\"hljs-string\">\"https://example.com/dynamic-content\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=<span class=\"hljs-string\">\"document.querySelector('.load-more-button').click();\"</span> <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\".content-item\"</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{result.extracted_content.count(<span class=\"hljs-string\">'.content-item'</span>)}</span> items\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(basic_session_crawl())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. Reusing the same <code>session_id</code> across multiple requests.\n2. Executing JavaScript to load more content dynamically.\n3. Properly closing the session to free resources.<p></p>\n<h2 id=\"advanced-technique-1-custom-execution-hooks\">Advanced Technique 1: Custom Execution Hooks</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">advanced_session_crawl_with_hooks</span>():\n    first_commit = <span class=\"hljs-string\">\"\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_execution_started</span>(<span class=\"hljs-params\">page</span>):\n        <span class=\"hljs-keyword\">nonlocal</span> first_commit\n        <span class=\"hljs-keyword\">try</span>:\n            <span class=\"hljs-keyword\">while</span> <span class=\"hljs-literal\">True</span>:\n                <span class=\"hljs-keyword\">await</span> page.wait_for_selector(<span class=\"hljs-string\">\"li.commit-item h4\"</span>)\n                commit = <span class=\"hljs-keyword\">await</span> page.query_selector(<span class=\"hljs-string\">\"li.commit-item h4\"</span>)\n                commit = <span class=\"hljs-keyword\">await</span> commit.evaluate(<span class=\"hljs-string\">\"(element) =&gt; element.textContent\"</span>).strip()\n                <span class=\"hljs-keyword\">if</span> commit <span class=\"hljs-keyword\">and</span> commit != first_commit:\n                    first_commit = commit\n                    <span class=\"hljs-keyword\">break</span>\n                <span class=\"hljs-keyword\">await</span> asyncio.sleep(<span class=\"hljs-number\">0.5</span>)\n        <span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Warning: New content didn't appear: <span class=\"hljs-subst\">{e}</span>\"</span>)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"commit_session\"</span>\n        url = <span class=\"hljs-string\">\"https://github.com/example/repo/commits/main\"</span>\n        crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"on_execution_started\"</span>, on_execution_started)\n\n        js_next_page = <span class=\"hljs-string\">\"\"\"document.querySelector('a.pagination-next').click();\"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=js_next_page <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\"li.commit-item\"</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.extracted_content)}</span> commits\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(advanced_session_crawl_with_hooks())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"advanced-technique-2-integrated-javascript-execution-and-waiting\">Advanced Technique 2: Integrated JavaScript Execution and Waiting</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">integrated_js_and_wait_crawl</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        session_id = <span class=\"hljs-string\">\"integrated_session\"</span>\n        url = <span class=\"hljs-string\">\"https://github.com/example/repo/commits/main\"</span>\n\n        js_next_page_and_wait = <span class=\"hljs-string\">\"\"\"\n        (async () =&gt; {\n            const getCurrentCommit = () =&gt; document.querySelector('li.commit-item h4').textContent.trim();\n            const initialCommit = getCurrentCommit();\n            document.querySelector('a.pagination-next').click();\n            while (getCurrentCommit() === initialCommit) {\n                await new Promise(resolve =&gt; setTimeout(resolve, 100));\n            }\n        })();\n        \"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">3</span>):\n            config = CrawlerRunConfig(\n                url=url,\n                session_id=session_id,\n                js_code=js_next_page_and_wait <span class=\"hljs-keyword\">if</span> page &gt; <span class=\"hljs-number\">0</span> <span class=\"hljs-keyword\">else</span> <span class=\"hljs-literal\">None</span>,\n                css_selector=<span class=\"hljs-string\">\"li.commit-item\"</span>,\n                js_only=page &gt; <span class=\"hljs-number\">0</span>,\n                cache_mode=CacheMode.BYPASS\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(config=config)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page + <span class=\"hljs-number\">1</span>}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.extracted_content)}</span> commits\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\nasyncio.run(integrated_js_and_wait_crawl())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n1. <strong>Authentication Flows</strong>: Login and interact with secured pages.\n2. <strong>Pagination Handling</strong>: Navigate through multiple pages.\n3. <strong>Form Submissions</strong>: Fill forms, submit, and process results.\n4. <strong>Multi-step Processes</strong>: Complete workflows that span multiple actions.<p></p>\n<h1 id=\"hooks-auth-in-asyncwebcrawler\">Hooks &amp; Auth in AsyncWebCrawler</h1>\n<p>1. <strong><code>on_browser_created</code></strong> – After browser creation.<br>\n2. <strong><code>on_page_context_created</code></strong> – After a new context &amp; page are created.<br>\n3. <strong><code>before_goto</code></strong> – Just before navigating to a page.<br>\n4. <strong><code>after_goto</code></strong> – Right after navigation completes.<br>\n5. <strong><code>on_user_agent_updated</code></strong> – Whenever the user agent changes.<br>\n6. <strong><code>on_execution_started</code></strong> – Once custom JavaScript execution begins.<br>\n7. <strong><code>before_retrieve_html</code></strong> – Just before the crawler retrieves final HTML.<br>\n8. <strong><code>before_return_html</code></strong> – Right before returning the HTML content.\n<strong>Important</strong>: Avoid heavy tasks in <code>on_browser_created</code> since you don’t yet have a page context. If you need to <em>log in</em>, do so in <strong><code>on_page_context_created</code></strong>.</p>\n<blockquote>\n<p>note \"Important Hook Usage Warning\"\n    <strong>Avoid Misusing Hooks</strong>: Do not manipulate page objects in the wrong hook or at the wrong time, as it can crash the pipeline or produce incorrect results. A common mistake is attempting to handle authentication prematurely—such as creating or closing pages in <code>on_browser_created</code>. \n  <strong>Use the Right Hook for Auth</strong>: If you need to log in or set tokens, use <code>on_page_context_created</code>. This ensures you have a valid page/context to work with, without disrupting the main crawling flow.</p>\n</blockquote>\n<h2 id=\"example-using-hooks-in-asyncwebcrawler\">Example: Using Hooks in AsyncWebCrawler</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> playwright.async_api <span class=\"hljs-keyword\">import</span> Page, BrowserContext\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"🔗 Hooks Example: Demonstrating recommended usage\"</span>)\n\n    <span class=\"hljs-comment\"># 1) Configure the browser</span>\n    browser_config = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># 2) Configure the crawler run</span>\n    crawler_run_config = CrawlerRunConfig(\n        js_code=<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        wait_for=<span class=\"hljs-string\">\"body\"</span>,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-comment\"># 3) Create the crawler instance</span>\n    crawler = AsyncWebCrawler(config=browser_config)\n\n    <span class=\"hljs-comment\">#</span>\n    <span class=\"hljs-comment\"># Define Hook Functions</span>\n    <span class=\"hljs-comment\">#</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_browser_created</span>(<span class=\"hljs-params\">browser, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called once the browser instance is created (but no pages or contexts yet)</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_browser_created - Browser created successfully!\"</span>)\n        <span class=\"hljs-comment\"># Typically, do minimal setup here if needed</span>\n        <span class=\"hljs-keyword\">return</span> browser\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_page_context_created</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called right after a new page + context are created (ideal for auth or route config).</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_page_context_created - Setting up page &amp; context.\"</span>)\n\n        <span class=\"hljs-comment\"># Example 1: Route filtering (e.g., block images)</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">route_filter</span>(<span class=\"hljs-params\">route</span>):\n            <span class=\"hljs-keyword\">if</span> route.request.resource_type == <span class=\"hljs-string\">\"image\"</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] Blocking image request: <span class=\"hljs-subst\">{route.request.url}</span>\"</span>)\n                <span class=\"hljs-keyword\">await</span> route.abort()\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-keyword\">await</span> route.continue_()\n\n        <span class=\"hljs-keyword\">await</span> context.route(<span class=\"hljs-string\">\"**\"</span>, route_filter)\n\n        <span class=\"hljs-comment\"># Example 2: (Optional) Simulate a login scenario</span>\n        <span class=\"hljs-comment\"># (We do NOT create or close pages here, just do quick steps if needed)</span>\n        <span class=\"hljs-comment\"># e.g., await page.goto(\"https://example.com/login\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.fill(\"input[name='username']\", \"testuser\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.fill(\"input[name='password']\", \"password123\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.click(\"button[type='submit']\")</span>\n        <span class=\"hljs-comment\"># e.g., await page.wait_for_selector(\"#welcome\")</span>\n        <span class=\"hljs-comment\"># e.g., await context.add_cookies([...])</span>\n        <span class=\"hljs-comment\"># Then continue</span>\n\n        <span class=\"hljs-comment\"># Example 3: Adjust the viewport</span>\n        <span class=\"hljs-keyword\">await</span> page.set_viewport_size({<span class=\"hljs-string\">\"width\"</span>: <span class=\"hljs-number\">1080</span>, <span class=\"hljs-string\">\"height\"</span>: <span class=\"hljs-number\">600</span>})\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_goto</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, url: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called before navigating to each URL.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] before_goto - About to navigate: <span class=\"hljs-subst\">{url}</span>\"</span>)\n        <span class=\"hljs-comment\"># e.g., inject custom headers</span>\n        <span class=\"hljs-keyword\">await</span> page.set_extra_http_headers({\n            <span class=\"hljs-string\">\"Custom-Header\"</span>: <span class=\"hljs-string\">\"my-value\"</span>\n        })\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">after_goto</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, \n        url: <span class=\"hljs-built_in\">str</span>, response, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called after navigation completes.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] after_goto - Successfully loaded: <span class=\"hljs-subst\">{url}</span>\"</span>)\n        <span class=\"hljs-comment\"># e.g., wait for a certain element if we want to verify</span>\n        <span class=\"hljs-keyword\">try</span>:\n            <span class=\"hljs-keyword\">await</span> page.wait_for_selector(<span class=\"hljs-string\">'.content'</span>, timeout=<span class=\"hljs-number\">1000</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] Found .content element!\"</span>)\n        <span class=\"hljs-keyword\">except</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] .content not found, continuing anyway.\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_user_agent_updated</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, \n        user_agent: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called whenever the user agent updates.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] on_user_agent_updated - New user agent: <span class=\"hljs-subst\">{user_agent}</span>\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">on_execution_started</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called after custom JavaScript execution begins.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] on_execution_started - JS code is running!\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_retrieve_html</span>(<span class=\"hljs-params\">page: Page, context: BrowserContext, **kwargs</span>):\n        <span class=\"hljs-comment\"># Called before final HTML retrieval.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[HOOK] before_retrieve_html - We can do final actions\"</span>)\n        <span class=\"hljs-comment\"># Example: Scroll again</span>\n        <span class=\"hljs-keyword\">await</span> page.evaluate(<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">before_return_html</span>(<span class=\"hljs-params\">\n        page: Page, context: BrowserContext, html: <span class=\"hljs-built_in\">str</span>, **kwargs\n    </span>):\n        <span class=\"hljs-comment\"># Called just before returning the HTML in the result.</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[HOOK] before_return_html - HTML length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(html)}</span>\"</span>)\n        <span class=\"hljs-keyword\">return</span> page\n\n    <span class=\"hljs-comment\">#</span>\n    <span class=\"hljs-comment\"># Attach Hooks</span>\n    <span class=\"hljs-comment\">#</span>\n\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"on_browser_created\"</span>, on_browser_created)\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_page_context_created\"</span>, on_page_context_created\n    )\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"before_goto\"</span>, before_goto)\n    crawler.crawler_strategy.set_hook(<span class=\"hljs-string\">\"after_goto\"</span>, after_goto)\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_user_agent_updated\"</span>, on_user_agent_updated\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"on_execution_started\"</span>, on_execution_started\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"before_retrieve_html\"</span>, before_retrieve_html\n    )\n    crawler.crawler_strategy.set_hook(\n        <span class=\"hljs-string\">\"before_return_html\"</span>, before_return_html\n    )\n\n    <span class=\"hljs-keyword\">await</span> crawler.start()\n\n    <span class=\"hljs-comment\"># 4) Run the crawler on an example page</span>\n    url = <span class=\"hljs-string\">\"https://example.com\"</span>\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config=crawler_run_config)\n\n    <span class=\"hljs-keyword\">if</span> result.success:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nCrawled URL:\"</span>, result.url)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.html))\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n    <span class=\"hljs-keyword\">await</span> crawler.close()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"hook-lifecycle-summary\">Hook Lifecycle Summary</h2>\n<p>1. <strong><code>on_browser_created</code></strong>:<br>\n   - Browser is up, but <strong>no</strong> pages or contexts yet.<br>\n   - Light setup only—don’t try to open or close pages here (that belongs in <code>on_page_context_created</code>).\n2. <strong><code>on_page_context_created</code></strong>:<br>\n   - Perfect for advanced <strong>auth</strong> or route blocking.<br>\n3. <strong><code>before_goto</code></strong>:<br>\n4. <strong><code>after_goto</code></strong>:<br>\n5. <strong><code>on_user_agent_updated</code></strong>:<br>\n   - Whenever the user agent changes (for stealth or different UA modes).\n6. <strong><code>on_execution_started</code></strong>:<br>\n   - If you set <code>js_code</code> or run custom scripts, this runs once your JS is about to start.\n7. <strong><code>before_retrieve_html</code></strong>:<br>\n8. <strong><code>before_return_html</code></strong>:<br>\n   - The last hook before returning HTML to the <code>CrawlResult</code>. Good for logging HTML length or minor modifications.</p>\n<h2 id=\"when-to-handle-authentication\">When to Handle Authentication</h2>\n<p><strong>Recommended</strong>: Use <strong><code>on_page_context_created</code></strong> if you need to:\n- Navigate to a login page or fill forms\n- Set cookies or localStorage tokens\n- Block resource routes to avoid ads\nThis ensures the newly created context is under your control <strong>before</strong> <code>arun()</code> navigates to the main URL.</p>\n<h2 id=\"additional-considerations\">Additional Considerations</h2>\n<ul>\n<li><strong>Session Management</strong>: If you want multiple <code>arun()</code> calls to reuse a single session, pass <code>session_id=</code> in your <code>CrawlerRunConfig</code>. Hooks remain the same.  </li>\n<li><strong>Concurrency</strong>: If you run <code>arun_many()</code>, each URL triggers these hooks in parallel. Ensure your hooks are thread/async-safe.</li>\n</ul>\n<h2 id=\"conclusion_1\">Conclusion</h2>\n<ul>\n<li><strong>Browser</strong> creation (light tasks only)</li>\n<li><strong>Page</strong> and <strong>context</strong> creation (auth, route blocking)</li>\n<li><strong>Navigation</strong> phases</li>\n<li><strong>Final HTML</strong> retrieval</li>\n<li><strong>Login</strong> or advanced tasks in <code>on_page_context_created</code>  </li>\n<li><strong>Custom headers</strong> or logs in <code>before_goto</code> / <code>after_goto</code>  </li>\n<li><strong>Scrolling</strong> or final checks in <code>before_retrieve_html</code> / <code>before_return_html</code></li>\n</ul>\n<hr>\n<h1 id=\"quick-reference\">Quick Reference</h1>\n<h2 id=\"core-imports\">Core Imports</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-javascript\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> <span class=\"hljs-title class_\">AsyncWebCrawler</span>, <span class=\"hljs-title class_\">BrowserConfig</span>, <span class=\"hljs-title class_\">CrawlerRunConfig</span>, <span class=\"hljs-title class_\">CacheMode</span>, <span class=\"hljs-title class_\">LLMConfig</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai.<span class=\"hljs-property\">extraction_strategy</span> <span class=\"hljs-keyword\">import</span> <span class=\"hljs-title class_\">LLMExtractionStrategy</span>, <span class=\"hljs-title class_\">JsonCssExtractionStrategy</span>, <span class=\"hljs-title class_\">CosineStrategy</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"basic-pattern\">Basic Pattern</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>)\n    print(result.markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"advanced-pattern\">Advanced Pattern</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, viewport_width=<span class=\"hljs-number\">1920</span>)\ncrawler_config = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS,\n    wait_for=<span class=\"hljs-string\">\"css:.content\"</span>,\n    screenshot=<span class=\"hljs-literal\">True</span>,\n    pdf=<span class=\"hljs-literal\">True</span>\n)\nstrategy = LLMExtractionStrategy(\n    llm_config=LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4\"</span>, api_token=<span class=\"hljs-string\">\"your-openai-token\"</span>),\n    instruction=<span class=\"hljs-string\">\"Extract products with name and price\"</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        config=crawler_config,\n        extraction_strategy=strategy\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"multi-url-pattern\">Multi-URL Pattern</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">urls</span> = [<span class=\"hljs-string\">\"https://example.com/1\"</span>, <span class=\"hljs-string\">\"https://example.com/2\"</span>]\n<span class=\"hljs-attr\">results</span> = await crawler.arun_many(urls, config=crawler_config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<p><strong>End of Crawl4AI SDK Documentation</strong></p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/adaptive-crawling/",
    "title": "Adaptive Crawling - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"adaptive-web-crawling\">Adaptive Web Crawling</h1>\n<h2 id=\"introduction\">Introduction</h2>\n<p>Traditional web crawlers follow predetermined patterns, crawling pages blindly without knowing when they've gathered enough information. <strong>Adaptive Crawling</strong> changes this paradigm by introducing intelligence into the crawling process.</p>\n<p>Think of it like research: when you're looking for information, you don't read every book in the library. You stop when you've found sufficient information to answer your question. That's exactly what Adaptive Crawling does for web scraping.</p>\n<h2 id=\"key-concepts\">Key Concepts</h2>\n<h3 id=\"the-problem-it-solves\">The Problem It Solves</h3>\n<p>When crawling websites for specific information, you face two challenges:\n1. <strong>Under-crawling</strong>: Stopping too early and missing crucial information\n2. <strong>Over-crawling</strong>: Wasting resources by crawling irrelevant pages</p>\n<p>Adaptive Crawling solves both by using a three-layer scoring system that determines when you have \"enough\" information.</p>\n<h3 id=\"how-it-works\">How It Works</h3>\n<p>The AdaptiveCrawler uses three metrics to measure information sufficiency:</p>\n<ul>\n<li><strong>Coverage</strong>: How well your collected pages cover the query terms</li>\n<li><strong>Consistency</strong>: Whether the information is coherent across pages  </li>\n<li><strong>Saturation</strong>: Detecting when new pages aren't adding new information</li>\n</ul>\n<p>When these metrics indicate sufficient information has been gathered, crawling stops automatically.</p>\n<h2 id=\"quick-start\">Quick Start</h2>\n<h3 id=\"basic-usage\">Basic Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, AdaptiveCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Create an adaptive crawler (config is optional)</span>\n        adaptive = AdaptiveCrawler(crawler)\n\n        <span class=\"hljs-comment\"># Start crawling with a query</span>\n        result = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n            start_url=<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>,\n            query=<span class=\"hljs-string\">\"async context managers\"</span>\n        )\n\n        <span class=\"hljs-comment\"># View statistics</span>\n        adaptive.print_stats()\n\n        <span class=\"hljs-comment\"># Get the most relevant content</span>\n        relevant_pages = adaptive.get_relevant_content(top_k=<span class=\"hljs-number\">5</span>)\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> relevant_pages:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"- <span class=\"hljs-subst\">{page[<span class=\"hljs-string\">'url'</span>]}</span> (score: <span class=\"hljs-subst\">{page[<span class=\"hljs-string\">'score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>)\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"configuration-options\">Configuration Options</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-vbnet\"><span class=\"hljs-keyword\">from</span> crawl4ai import AdaptiveConfig\n\nconfig = AdaptiveConfig(\n    confidence_threshold=<span class=\"hljs-number\">0.8</span>,    # <span class=\"hljs-keyword\">Stop</span> <span class=\"hljs-keyword\">when</span> <span class=\"hljs-number\">80%</span> confident (<span class=\"hljs-keyword\">default</span>: <span class=\"hljs-number\">0.7</span>)\n    max_pages=<span class=\"hljs-number\">30</span>,               # Maximum pages <span class=\"hljs-keyword\">to</span> crawl (<span class=\"hljs-keyword\">default</span>: <span class=\"hljs-number\">20</span>)\n    top_k_links=<span class=\"hljs-number\">5</span>,              # Links <span class=\"hljs-keyword\">to</span> follow per page (<span class=\"hljs-keyword\">default</span>: <span class=\"hljs-number\">3</span>)\n    min_gain_threshold=<span class=\"hljs-number\">0.05</span>     # Minimum expected gain <span class=\"hljs-keyword\">to</span> <span class=\"hljs-keyword\">continue</span> (<span class=\"hljs-keyword\">default</span>: <span class=\"hljs-number\">0.1</span>)\n)\n\nadaptive = AdaptiveCrawler(crawler, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"crawling-strategies\">Crawling Strategies</h2>\n<p>Adaptive Crawling supports two distinct strategies for determining information sufficiency:</p>\n<h3 id=\"statistical-strategy-default\">Statistical Strategy (Default)</h3>\n<p>The statistical strategy uses pure information theory and term-based analysis:</p>\n<ul>\n<li><strong>Fast and efficient</strong> - No API calls or model loading</li>\n<li><strong>Term-based coverage</strong> - Analyzes query term presence and distribution</li>\n<li><strong>No external dependencies</strong> - Works offline</li>\n<li><strong>Best for</strong>: Well-defined queries with specific terminology</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Default configuration uses statistical strategy</span>\nconfig = AdaptiveConfig(\n    strategy=<span class=\"hljs-string\">\"statistical\"</span>,  <span class=\"hljs-comment\"># This is the default</span>\n    confidence_threshold=0.8\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"embedding-strategy\">Embedding Strategy</h3>\n<p>The embedding strategy uses semantic embeddings for deeper understanding:</p>\n<ul>\n<li><strong>Semantic understanding</strong> - Captures meaning beyond exact term matches</li>\n<li><strong>Query expansion</strong> - Automatically generates query variations</li>\n<li><strong>Gap-driven selection</strong> - Identifies semantic gaps in knowledge</li>\n<li><strong>Validation-based stopping</strong> - Uses held-out queries to validate coverage</li>\n<li><strong>Best for</strong>: Complex queries, ambiguous topics, conceptual understanding</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Configure embedding strategy with local embeddings</span>\nconfig = AdaptiveConfig(\n    strategy=<span class=\"hljs-string\">\"embedding\"</span>,\n    embedding_model=<span class=\"hljs-string\">\"sentence-transformers/all-MiniLM-L6-v2\"</span>,  <span class=\"hljs-comment\"># Default</span>\n    n_query_variations=<span class=\"hljs-number\">10</span>,  <span class=\"hljs-comment\"># Generate 10 query variations</span>\n    embedding_min_confidence_threshold=<span class=\"hljs-number\">0.1</span>  <span class=\"hljs-comment\"># Stop if completely irrelevant</span>\n)\n\n<span class=\"hljs-comment\"># With separate LLM configs for embeddings and query expansion (recommended)</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\nconfig = AdaptiveConfig(\n    strategy=<span class=\"hljs-string\">\"embedding\"</span>,\n    <span class=\"hljs-comment\"># Embedding model — used for text-to-vector calls</span>\n    embedding_llm_config=LLMConfig(\n        provider=<span class=\"hljs-string\">'openai/text-embedding-3-small'</span>,\n        api_token=<span class=\"hljs-string\">'your-api-key'</span>\n    ),\n    <span class=\"hljs-comment\"># Query model — used for chat completion (query expansion)</span>\n    query_llm_config=LLMConfig(\n        provider=<span class=\"hljs-string\">'openai/gpt-4o-mini'</span>,\n        api_token=<span class=\"hljs-string\">'your-api-key'</span>\n    )\n)\n\n<span class=\"hljs-comment\"># Alternative: Dictionary format (backward compatible)</span>\nconfig = AdaptiveConfig(\n    strategy=<span class=\"hljs-string\">\"embedding\"</span>,\n    embedding_llm_config={\n        <span class=\"hljs-string\">'provider'</span>: <span class=\"hljs-string\">'openai/text-embedding-3-small'</span>,\n        <span class=\"hljs-string\">'api_token'</span>: <span class=\"hljs-string\">'your-api-key'</span>\n    },\n    query_llm_config={\n        <span class=\"hljs-string\">'provider'</span>: <span class=\"hljs-string\">'openai/gpt-4o-mini'</span>,\n        <span class=\"hljs-string\">'api_token'</span>: <span class=\"hljs-string\">'your-api-key'</span>\n    }\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<blockquote>\n<p><strong>Note:</strong> The embedding strategy makes two types of API calls that need different model types:\n- <strong>Embedding calls</strong> (text → vector) require an embedding model like <code>text-embedding-3-small</code>\n- <strong>Query expansion</strong> (chat completion) requires a chat model like <code>gpt-4o-mini</code></p>\n<p>Use <code>embedding_llm_config</code> for the embedding model and <code>query_llm_config</code> for the chat model. If <code>query_llm_config</code> is not set, it falls back to <code>embedding_llm_config</code> for backward compatibility.</p>\n</blockquote>\n<h3 id=\"strategy-comparison\">Strategy Comparison</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th>Statistical</th>\n<th>Embedding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Speed</strong></td>\n<td>Very fast</td>\n<td>Moderate (API calls)</td>\n</tr>\n<tr>\n<td><strong>Cost</strong></td>\n<td>Free</td>\n<td>Depends on provider</td>\n</tr>\n<tr>\n<td><strong>Accuracy</strong></td>\n<td>Good for exact terms</td>\n<td>Excellent for concepts</td>\n</tr>\n<tr>\n<td><strong>Dependencies</strong></td>\n<td>None</td>\n<td>Embedding model/API</td>\n</tr>\n<tr>\n<td><strong>Query Understanding</strong></td>\n<td>Literal</td>\n<td>Semantic</td>\n</tr>\n<tr>\n<td><strong>Best Use Case</strong></td>\n<td>Technical docs, specific terms</td>\n<td>Research, broad topics</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"embedding-strategy-configuration\">Embedding Strategy Configuration</h3>\n<p>The embedding strategy offers fine-tuned control through several parameters:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = AdaptiveConfig(\n    strategy=<span class=\"hljs-string\">\"embedding\"</span>,\n\n    <span class=\"hljs-comment\"># Model configuration</span>\n    embedding_model=<span class=\"hljs-string\">\"sentence-transformers/all-MiniLM-L6-v2\"</span>,\n    embedding_llm_config=<span class=\"hljs-literal\">None</span>,  <span class=\"hljs-comment\"># Use for API-based embeddings (embedding model)</span>\n    query_llm_config=<span class=\"hljs-literal\">None</span>,  <span class=\"hljs-comment\"># Use for query expansion (chat completion model)</span>\n\n    <span class=\"hljs-comment\"># Query expansion</span>\n    n_query_variations=<span class=\"hljs-number\">10</span>,  <span class=\"hljs-comment\"># Number of query variations to generate</span>\n\n    <span class=\"hljs-comment\"># Coverage parameters</span>\n    embedding_coverage_radius=<span class=\"hljs-number\">0.2</span>,  <span class=\"hljs-comment\"># Distance threshold for coverage</span>\n    embedding_k_exp=<span class=\"hljs-number\">3.0</span>,  <span class=\"hljs-comment\"># Exponential decay factor (higher = stricter)</span>\n\n    <span class=\"hljs-comment\"># Stopping criteria</span>\n    embedding_min_relative_improvement=<span class=\"hljs-number\">0.1</span>,  <span class=\"hljs-comment\"># Min improvement to continue</span>\n    embedding_validation_min_score=<span class=\"hljs-number\">0.3</span>,  <span class=\"hljs-comment\"># Min validation score</span>\n    embedding_min_confidence_threshold=<span class=\"hljs-number\">0.1</span>,  <span class=\"hljs-comment\"># Below this = irrelevant</span>\n\n    <span class=\"hljs-comment\"># Link selection</span>\n    embedding_overlap_threshold=<span class=\"hljs-number\">0.85</span>,  <span class=\"hljs-comment\"># Similarity for deduplication</span>\n\n    <span class=\"hljs-comment\"># Display confidence mapping</span>\n    embedding_quality_min_confidence=<span class=\"hljs-number\">0.7</span>,  <span class=\"hljs-comment\"># Min displayed confidence</span>\n    embedding_quality_max_confidence=<span class=\"hljs-number\">0.95</span>  <span class=\"hljs-comment\"># Max displayed confidence</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"handling-irrelevant-queries\">Handling Irrelevant Queries</h3>\n<p>The embedding strategy can detect when a query is completely unrelated to the content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># This will stop quickly with low confidence</span>\nresult = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n    start_url=<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>,\n    query=<span class=\"hljs-string\">\"how to cook pasta\"</span>  <span class=\"hljs-comment\"># Irrelevant to Python docs</span>\n)\n\n<span class=\"hljs-comment\"># Check if query was irrelevant</span>\n<span class=\"hljs-keyword\">if</span> result.metrics.get(<span class=\"hljs-string\">'is_irrelevant'</span>, <span class=\"hljs-literal\">False</span>):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Query is unrelated to the content!\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"when-to-use-adaptive-crawling\">When to Use Adaptive Crawling</h2>\n<h3 id=\"perfect-for\">Perfect For:</h3>\n<ul>\n<li><strong>Research Tasks</strong>: Finding comprehensive information about a topic</li>\n<li><strong>Question Answering</strong>: Gathering sufficient context to answer specific queries</li>\n<li><strong>Knowledge Base Building</strong>: Creating focused datasets for AI/ML applications</li>\n<li><strong>Competitive Intelligence</strong>: Collecting complete information about specific products/features</li>\n</ul>\n<h3 id=\"not-recommended-for\">Not Recommended For:</h3>\n<ul>\n<li><strong>Full Site Archiving</strong>: When you need every page regardless of content</li>\n<li><strong>Structured Data Extraction</strong>: When targeting specific, known page patterns</li>\n<li><strong>Real-time Monitoring</strong>: When you need continuous updates</li>\n</ul>\n<h2 id=\"understanding-the-output\">Understanding the Output</h2>\n<h3 id=\"confidence-score\">Confidence Score</h3>\n<p>The confidence score (0-1) indicates how sufficient the gathered information is:\n- <strong>0.0-0.3</strong>: Insufficient information, needs more crawling\n- <strong>0.3-0.6</strong>: Partial information, may answer basic queries\n- <strong>0.6-0.7</strong>: Good coverage, can answer most queries\n- <strong>0.7-1.0</strong>: Excellent coverage, comprehensive information</p>\n<h3 id=\"statistics-display\">Statistics Display</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">adaptive.print_stats<span class=\"hljs-punctuation\">(</span>detailed<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span><span class=\"hljs-punctuation\">)</span>  <span class=\"hljs-comment\"># Summary table</span>\nadaptive.print_stats<span class=\"hljs-punctuation\">(</span>detailed<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span><span class=\"hljs-punctuation\">)</span>   <span class=\"hljs-comment\"># Detailed metrics</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The summary shows:\n- Pages crawled vs. confidence achieved\n- Coverage, consistency, and saturation scores\n- Crawling efficiency metrics</p>\n<h2 id=\"persistence-and-resumption\">Persistence and Resumption</h2>\n<h3 id=\"saving-progress\">Saving Progress</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> AdaptiveConfig<span class=\"hljs-punctuation\">(</span>\n    save_state<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    state_path<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"my_crawl_state.json\"</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Crawl will auto-save progress</span>\nresult <span class=\"hljs-punctuation\">=</span> await adaptive.digest<span class=\"hljs-punctuation\">(</span>start_url, <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"resuming-a-crawl\">Resuming a Crawl</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Resume from saved state</span>\nresult <span class=\"hljs-punctuation\">=</span> await adaptive.digest<span class=\"hljs-punctuation\">(</span>\n    start_url,\n    <span class=\"hljs-keyword\">query</span>,\n    resume_from<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"my_crawl_state.json\"</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"exporting-knowledge-base\">Exporting Knowledge Base</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Export collected pages to JSONL</span>\nadaptive.export_knowledge_base(<span class=\"hljs-string\">\"knowledge_base.jsonl\"</span>)\n\n<span class=\"hljs-comment\"># Import into another session</span>\nnew_adaptive = AdaptiveCrawler(crawler)\nawait new_adaptive.import_knowledge_base(<span class=\"hljs-string\">\"knowledge_base.jsonl\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices\">Best Practices</h2>\n<h3 id=\"1-query-formulation\">1. Query Formulation</h3>\n<ul>\n<li>Use specific, descriptive queries</li>\n<li>Include key terms you expect to find</li>\n<li>Avoid overly broad queries</li>\n</ul>\n<h3 id=\"2-threshold-tuning\">2. Threshold Tuning</h3>\n<ul>\n<li>Start with default (0.7) for general use</li>\n<li>Lower to 0.5-0.6 for exploratory crawling</li>\n<li>Raise to 0.8+ for exhaustive coverage</li>\n</ul>\n<h3 id=\"3-performance-optimization\">3. Performance Optimization</h3>\n<ul>\n<li>Use appropriate <code>max_pages</code> limits</li>\n<li>Adjust <code>top_k_links</code> based on site structure</li>\n<li>Enable caching for repeat crawls</li>\n</ul>\n<h3 id=\"4-link-selection\">4. Link Selection</h3>\n<ul>\n<li>The crawler prioritizes links based on:</li>\n<li>Relevance to query</li>\n<li>Expected information gain</li>\n<li>URL structure and depth</li>\n</ul>\n<h2 id=\"examples\">Examples</h2>\n<h3 id=\"research-assistant\">Research Assistant</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Gather information about a programming concept</span>\nresult = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n    start_url=<span class=\"hljs-string\">\"https://realpython.com\"</span>,\n    query=<span class=\"hljs-string\">\"python decorators implementation patterns\"</span>\n)\n\n<span class=\"hljs-comment\"># Get the most relevant excerpts</span>\n<span class=\"hljs-keyword\">for</span> doc <span class=\"hljs-keyword\">in</span> adaptive.get_relevant_content(top_k=<span class=\"hljs-number\">3</span>):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\nFrom: <span class=\"hljs-subst\">{doc[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Relevance: <span class=\"hljs-subst\">{doc[<span class=\"hljs-string\">'score'</span>]:<span class=\"hljs-number\">.2</span>%}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(doc[<span class=\"hljs-string\">'content'</span>][:<span class=\"hljs-number\">500</span>] + <span class=\"hljs-string\">\"...\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"knowledge-base-builder\">Knowledge Base Builder</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Build a focused knowledge base about machine learning</span>\nqueries <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-string\">\"supervised learning algorithms\"</span>,\n    <span class=\"hljs-string\">\"neural network architectures\"</span>, \n    <span class=\"hljs-string\">\"model evaluation metrics\"</span>\n<span class=\"hljs-punctuation\">]</span>\n\nfor <span class=\"hljs-keyword\">query</span> in <span class=\"hljs-symbol\">queries</span><span class=\"hljs-punctuation\">:</span>\n    await adaptive.digest<span class=\"hljs-punctuation\">(</span>\n        start_url<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"https://scikit-learn.org/stable/\"</span>,\n        <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-keyword\">query</span>\n    <span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Export combined knowledge base</span>\nadaptive.export_knowledge_base<span class=\"hljs-punctuation\">(</span><span class=\"hljs-string\">\"ml_knowledge.jsonl\"</span><span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"api-documentation-crawler\">API Documentation Crawler</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Intelligently crawl API documentation</span>\nconfig = AdaptiveConfig(\n    confidence_threshold=0.85,  <span class=\"hljs-comment\"># Higher threshold for completeness</span>\n    max_pages=30\n)\n\nadaptive = AdaptiveCrawler(crawler, config)\nresult = await adaptive.digest(\n    start_url=<span class=\"hljs-string\">\"https://api.example.com/docs\"</span>,\n    query=<span class=\"hljs-string\">\"authentication endpoints rate limits\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"next-steps\">Next Steps</h2>\n<ul>\n<li>Learn about <a href=\"../../advanced/adaptive-strategies/\">Advanced Adaptive Strategies</a></li>\n<li>Explore the <a href=\"../../api/adaptive-crawler/\">AdaptiveCrawler API Reference</a></li>\n<li>See more <a href=\"https://github.com/unclecode/crawl4ai/tree/main/docs/examples/adaptive_crawling\">Examples</a></li>\n</ul>\n<h2 id=\"faq\">FAQ</h2>\n<p><strong>Q: How is this different from traditional crawling?</strong>\nA: Traditional crawling follows fixed patterns (BFS/DFS). Adaptive crawling makes intelligent decisions about which links to follow and when to stop based on information gain.</p>\n<p><strong>Q: Can I use this with JavaScript-heavy sites?</strong>\nA: Yes! AdaptiveCrawler inherits all capabilities from AsyncWebCrawler, including JavaScript execution.</p>\n<p><strong>Q: How does it handle large websites?</strong>\nA: The algorithm naturally limits crawling to relevant sections. Use <code>max_pages</code> as a safety limit.</p>\n<p><strong>Q: Can I customize the scoring algorithms?</strong>\nA: Advanced users can implement custom strategies. See <a href=\"../../advanced/adaptive-strategies/\">Adaptive Strategies</a>.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/ask-ai/",
    "title": "Ask AI - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <div class=\"ask-ai-container\">\n\n</div>\n\n\n\n\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/browser-crawler-config/",
    "title": "Browser, Crawler & LLM Config - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"browser-crawler-llm-configuration-quick-overview\">Browser, Crawler &amp; LLM Configuration (Quick Overview)</h1>\n<p>Crawl4AI's flexibility stems from two key classes:</p>\n<ol>\n<li><strong><code>BrowserConfig</code></strong> – Dictates <strong>how</strong> the browser is launched and behaves (e.g., headless or visible, proxy, user agent).  </li>\n<li><strong><code>CrawlerRunConfig</code></strong> – Dictates <strong>how</strong> each <strong>crawl</strong> operates (e.g., caching, extraction, timeouts, JavaScript code to run, etc.).  </li>\n<li><strong><code>LLMConfig</code></strong> - Dictates <strong>how</strong> LLM providers are configured. (model, api token, base url, temperature etc.)</li>\n</ol>\n<p>In most examples, you create <strong>one</strong> <code>BrowserConfig</code> for the entire crawler session, then pass a <strong>fresh</strong> or re-used <code>CrawlerRunConfig</code> whenever you call <code>arun()</code>. This tutorial shows the most commonly used parameters. If you need advanced or rarely used fields, see the <a href=\"../../api/parameters/\">Configuration Parameters</a>.</p>\n<hr>\n<h2 id=\"1-browserconfig-essentials\">1. BrowserConfig Essentials</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">BrowserConfig</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        browser_type=<span class=\"hljs-string\">\"chromium\"</span>,\n        headless=<span class=\"hljs-literal\">True</span>,\n        browser_mode=<span class=\"hljs-string\">\"dedicated\"</span>,\n        use_managed_browser=<span class=\"hljs-literal\">False</span>,\n        cdp_url=<span class=\"hljs-literal\">None</span>,\n        debugging_port=<span class=\"hljs-number\">9222</span>,\n        host=<span class=\"hljs-string\">\"localhost\"</span>,\n        proxy_config=<span class=\"hljs-literal\">None</span>,\n        viewport_width=<span class=\"hljs-number\">1080</span>,\n        viewport_height=<span class=\"hljs-number\">600</span>,\n        verbose=<span class=\"hljs-literal\">True</span>,\n        use_persistent_context=<span class=\"hljs-literal\">False</span>,\n        user_data_dir=<span class=\"hljs-literal\">None</span>,\n        cookies=<span class=\"hljs-literal\">None</span>,\n        headers=<span class=\"hljs-literal\">None</span>,\n        user_agent=(<span class=\"hljs-params\">\n            <span class=\"hljs-comment\"># \"Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:109.0) AppleWebKit/537.36 \"</span>\n            <span class=\"hljs-comment\"># \"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 \"</span>\n            <span class=\"hljs-comment\"># \"(KHTML, like Gecko) Chrome/116.0.5845.187 Safari/604.1 Edg/117.0.2045.47\"</span>\n            <span class=\"hljs-string\">\"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 Chrome/116.0.0.0 Safari/537.36\"</span>\n        </span>),\n        user_agent_mode=<span class=\"hljs-string\">\"\"</span>,\n        text_mode=<span class=\"hljs-literal\">False</span>,\n        light_mode=<span class=\"hljs-literal\">False</span>,\n        extra_args=<span class=\"hljs-literal\">None</span>,\n        enable_stealth=<span class=\"hljs-literal\">False</span>,\n        <span class=\"hljs-comment\"># ... other advanced parameters omitted here</span>\n    </span>):\n        ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"key-fields-to-note\">Key Fields to Note</h3>\n<p>1.⠀<strong><code>browser_type</code></strong><br>\n   - Options: <code>\"chromium\"</code>, <code>\"firefox\"</code>, or <code>\"webkit\"</code>.<br>\n   - Defaults to <code>\"chromium\"</code>.<br>\n   - If you need a different engine, specify it here.</p>\n<p>2.⠀<strong><code>headless</code></strong><br>\n   - <code>True</code>: Runs the browser in headless mode (invisible browser).<br>\n   - <code>False</code>: Runs the browser in visible mode, which helps with debugging.</p>\n<p>3.⠀<strong><code>browser_mode</code></strong><br>\n   - Determines how the browser should be initialized:\n     - <code>\"dedicated\"</code> (default): Creates a new browser instance each time\n     - <code>\"builtin\"</code>: Uses the builtin CDP browser running in background\n     - <code>\"custom\"</code>: Uses explicit CDP settings provided in <code>cdp_url</code>\n     - <code>\"docker\"</code>: Runs browser in Docker container with isolation</p>\n<p>4.⠀<strong><code>use_managed_browser</code></strong> &amp; <strong><code>cdp_url</code></strong><br>\n   - <code>use_managed_browser=True</code>: Launch browser using Chrome DevTools Protocol (CDP) for advanced control\n   - <code>cdp_url</code>: URL for CDP endpoint (e.g., <code>\"ws://localhost:9222/devtools/browser/\"</code>)\n   - Automatically set based on <code>browser_mode</code></p>\n<p>5.⠀<strong><code>debugging_port</code></strong> &amp; <strong><code>host</code></strong><br>\n   - <code>debugging_port</code>: Port for browser debugging protocol (default: 9222)\n   - <code>host</code>: Host for browser connection (default: \"localhost\")</p>\n<p>6.⠀<strong><code>proxy_config</code></strong><br>\n   - A <code>ProxyConfig</code> object or dictionary with fields like:<br>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-attr\">\"server\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"http://proxy.example.com:8080\"</span><span class=\"hljs-punctuation\">,</span> \n    <span class=\"hljs-attr\">\"username\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span><span class=\"hljs-punctuation\">,</span> \n    <span class=\"hljs-attr\">\"password\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"...\"</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n   - Leave as <code>None</code> if a proxy is not required.<p></p>\n<p>7.⠀<strong><code>viewport_width</code> &amp; <code>viewport_height</code></strong>\n   - The initial window size.\n   - Some sites behave differently with smaller or bigger viewports.</p>\n<p>8.⠀<strong><code>device_scale_factor</code></strong>\n   - Controls the device pixel ratio (DPR) for rendering. Default is <code>1.0</code>.\n   - Set to <code>2.0</code> for Retina-quality screenshots (e.g., a 1920×1080 viewport produces 3840×2160 images).\n   - Higher values increase screenshot size and rendering time proportionally.</p>\n<p>9.⠀<strong><code>verbose</code></strong><br>\n   - If <code>True</code>, prints extra logs.<br>\n   - Handy for debugging.</p>\n<p>9.⠀<strong><code>use_persistent_context</code></strong><br>\n   - If <code>True</code>, uses a <strong>persistent</strong> browser profile, storing cookies/local storage across runs.<br>\n   - Typically also set <code>user_data_dir</code> to point to a folder.</p>\n<p>10.⠀<strong><code>cookies</code></strong> &amp; <strong><code>headers</code></strong><br>\n    - If you want to start with specific cookies or add universal HTTP headers to the browser context, set them here.<br>\n    - E.g. <code>cookies=[{\"name\": \"session\", \"value\": \"abc123\", \"domain\": \"example.com\"}]</code>.</p>\n<p>11.⠀<strong><code>user_agent</code></strong> &amp; <strong><code>user_agent_mode</code></strong><br>\n    - <code>user_agent</code>: Custom User-Agent string. If <code>None</code>, a default is used.<br>\n    - <code>user_agent_mode</code>: Set to <code>\"random\"</code> for randomization (helps fight bot detection).</p>\n<p>12.⠀<strong><code>text_mode</code></strong> &amp; <strong><code>light_mode</code></strong>\n    - <code>text_mode=True</code> disables images, possibly speeding up text-only crawls.\n    - <code>light_mode=True</code> turns off certain background features for performance.</p>\n<p>13.⠀<strong><code>avoid_ads</code></strong> &amp; <strong><code>avoid_css</code></strong>\n    - <code>avoid_ads=True</code> blocks requests to common ad and tracker domains (Google Analytics, DoubleClick, Facebook, Hotjar, etc.) at the browser context level. Reduces network overhead and memory usage.\n    - <code>avoid_css=True</code> blocks loading of CSS files (<code>.css</code>, <code>.less</code>, <code>.scss</code>, <code>.sass</code>), useful when you only need text content and want faster, leaner crawls.\n    - Both default to <code>False</code> (opt-in). Can be combined with each other and with <code>text_mode</code>.</p>\n<p>14.⠀<strong><code>extra_args</code></strong><br>\n    - Additional flags for the underlying browser.<br>\n    - E.g. <code>[\"--disable-extensions\"]</code>.</p>\n<p>15.⠀<strong><code>enable_stealth</code></strong>\n    - If <code>True</code>, enables stealth mode using playwright-stealth.\n    - Modifies browser fingerprints to avoid basic bot detection.\n    - Default is <code>False</code>. Recommended for sites with bot protection.</p>\n<h3 id=\"helper-methods\">Helper Methods</h3>\n<p>Both configuration classes provide a <code>clone()</code> method to create modified copies:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Create a base browser config</span>\nbase_browser <span class=\"hljs-punctuation\">=</span> BrowserConfig<span class=\"hljs-punctuation\">(</span>\n    browser_type<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"chromium\"</span>,\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    text_mode<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Create a visible browser config for debugging</span>\ndebug_browser <span class=\"hljs-punctuation\">=</span> base_browser.clone<span class=\"hljs-punctuation\">(</span>\n    headless<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>,\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"class-level-defaults\">Class-Level Defaults</h3>\n<p>Both <code>BrowserConfig</code> and <code>CrawlerRunConfig</code> support <strong>class-level default overrides</strong> via <code>set_defaults()</code>. This is useful in server/cloud deployments where every config instance needs the same base settings — set them once at startup instead of repeating at every call site.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-vbnet\"><span class=\"hljs-keyword\">from</span> crawl4ai import BrowserConfig, CrawlerRunConfig\n\n# At application startup — one time\nBrowserConfig.set_defaults(\n    cache_cdp_connection=<span class=\"hljs-literal\">True</span>,\n    cdp_close_delay=<span class=\"hljs-number\">0</span>,\n    create_isolated_context=<span class=\"hljs-literal\">True</span>,\n)\nCrawlerRunConfig.set_defaults(verbose=<span class=\"hljs-literal\">False</span>)\n\n# Every <span class=\"hljs-built_in\">new</span> instance automatically <span class=\"hljs-keyword\">inherits</span> those defaults\ncfg = BrowserConfig(cdp_url=<span class=\"hljs-string\">\"ws://localhost:9222\"</span>)\n# → cache_cdp_connection=<span class=\"hljs-literal\">True</span>, cdp_close_delay=<span class=\"hljs-number\">0</span>, create_isolated_context=<span class=\"hljs-literal\">True</span>\n\n# <span class=\"hljs-keyword\">Explicit</span> values still win\ncfg = BrowserConfig(cdp_url=<span class=\"hljs-string\">\"ws://localhost:9222\"</span>, cache_cdp_connection=<span class=\"hljs-literal\">False</span>)\n# → cache_cdp_connection=<span class=\"hljs-literal\">False</span> (<span class=\"hljs-keyword\">explicit</span> <span class=\"hljs-keyword\">overrides</span> the <span class=\"hljs-keyword\">class</span> <span class=\"hljs-keyword\">default</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Available methods</strong> (on both <code>BrowserConfig</code> and <code>CrawlerRunConfig</code>):</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Method</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>set_defaults(**kwargs)</code></td>\n<td>Set class-level defaults. Invalid parameter names raise <code>ValueError</code>.</td>\n</tr>\n<tr>\n<td><code>get_defaults()</code></td>\n<td>Return a copy of the current class-level defaults.</td>\n</tr>\n<tr>\n<td><code>reset_defaults()</code></td>\n<td>Clear all class-level defaults.</td>\n</tr>\n<tr>\n<td><code>reset_defaults(\"param1\", \"param2\")</code></td>\n<td>Clear only the named defaults.</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n<p><strong>Note:</strong> Class defaults are independent per class — <code>BrowserConfig.set_defaults()</code> does not affect <code>CrawlerRunConfig</code>, and vice versa. Defaults are stored in memory and apply for the lifetime of the process.</p>\n</blockquote>\n<p><strong>Minimal Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig\n\nbrowser_conf = BrowserConfig(\n    browser_type=<span class=\"hljs-string\">\"firefox\"</span>,\n    headless=<span class=\"hljs-literal\">False</span>,\n    text_mode=<span class=\"hljs-literal\">True</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n    <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"2-crawlerrunconfig-essentials\">2. CrawlerRunConfig Essentials</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CrawlerRunConfig</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">\n        word_count_threshold=<span class=\"hljs-number\">200</span>,\n        extraction_strategy=<span class=\"hljs-literal\">None</span>,\n        chunking_strategy=RegexChunking(<span class=\"hljs-params\"></span>),\n        markdown_generator=<span class=\"hljs-literal\">None</span>,\n        cache_mode=CacheMode.BYPASS,\n        js_code=<span class=\"hljs-literal\">None</span>,\n        c4a_script=<span class=\"hljs-literal\">None</span>,\n        wait_for=<span class=\"hljs-literal\">None</span>,\n        screenshot=<span class=\"hljs-literal\">False</span>,\n        pdf=<span class=\"hljs-literal\">False</span>,\n        capture_mhtml=<span class=\"hljs-literal\">False</span>,\n        <span class=\"hljs-comment\"># Location and Identity Parameters</span>\n        locale=<span class=\"hljs-literal\">None</span>,            <span class=\"hljs-comment\"># e.g. \"en-US\", \"fr-FR\"</span>\n        timezone_id=<span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># e.g. \"America/New_York\"</span>\n        geolocation=<span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># GeolocationConfig object</span>\n        <span class=\"hljs-comment\"># Proxy Configuration</span>\n        proxy_config=<span class=\"hljs-literal\">None</span>,\n        proxy_rotation_strategy=<span class=\"hljs-literal\">None</span>,\n        <span class=\"hljs-comment\"># Page Interaction Parameters</span>\n        scan_full_page=<span class=\"hljs-literal\">False</span>,\n        scroll_delay=<span class=\"hljs-number\">0.2</span>,\n        wait_until=<span class=\"hljs-string\">\"domcontentloaded\"</span>,\n        page_timeout=<span class=\"hljs-number\">60000</span>,\n        delay_before_return_html=<span class=\"hljs-number\">0.1</span>,\n        <span class=\"hljs-comment\"># URL Matching Parameters</span>\n        url_matcher=<span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># For URL-specific configurations</span>\n        match_mode=MatchMode.OR,\n        verbose=<span class=\"hljs-literal\">True</span>,\n        stream=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Enable streaming for arun_many()</span>\n        <span class=\"hljs-comment\"># ... other advanced parameters omitted</span>\n    </span>):\n        ...\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"key-fields-to-note_1\">Key Fields to Note</h3>\n<p>1.⠀<strong><code>word_count_threshold</code></strong>:<br>\n   - The minimum word count before a block is considered.<br>\n   - If your site has lots of short paragraphs or items, you can lower it.</p>\n<p>2.⠀<strong><code>extraction_strategy</code></strong>:<br>\n   - Where you plug in JSON-based extraction (CSS, LLM, etc.).<br>\n   - If <code>None</code>, no structured extraction is done (only raw/cleaned HTML + markdown).</p>\n<p>3.⠀<strong><code>chunking_strategy</code></strong>:<br>\n   - Strategy to chunk content before extraction.<br>\n   - Defaults to <code>RegexChunking()</code>. Can be customized for different chunking approaches.</p>\n<p>4.⠀<strong><code>markdown_generator</code></strong>:<br>\n   - E.g., <code>DefaultMarkdownGenerator(...)</code>, controlling how HTML→Markdown conversion is done.<br>\n   - If <code>None</code>, a default approach is used.</p>\n<p>5.⠀<strong><code>cache_mode</code></strong>:<br>\n   - Controls caching behavior (<code>ENABLED</code>, <code>BYPASS</code>, <code>DISABLED</code>, etc.).<br>\n   - Defaults to <code>CacheMode.BYPASS</code>.</p>\n<p>6.⠀<strong><code>js_code</code></strong>, <strong><code>js_code_before_wait</code></strong>, &amp; <strong><code>c4a_script</code></strong>:\n   - <code>js_code</code>: JavaScript to run <strong>after</strong> <code>wait_for</code> completes — on the fully-loaded page.\n   - <code>js_code_before_wait</code>: JavaScript to run <strong>before</strong> <code>wait_for</code> — for triggering loading that <code>wait_for</code> then checks.\n   - <code>c4a_script</code>: C4A script that compiles to JavaScript.\n   - Great for \"Load More\" buttons or user interactions.</p>\n<p>7.⠀<strong><code>wait_for</code></strong>:<br>\n   - A CSS or JS expression to wait for before extracting content.<br>\n   - Common usage: <code>wait_for=\"css:.main-loaded\"</code> or <code>wait_for=\"js:() =&gt; window.loaded === true\"</code>.</p>\n<p>8.⠀<strong><code>flatten_shadow_dom</code></strong>:\n   - If <code>True</code>, flattens Shadow DOM content into the light DOM before HTML capture.\n   - Essential for sites built with Web Components (Stencil, Lit, Shoelace, etc.).\n   - Also force-opens closed shadow roots. See <a href=\"../content-selection/#31-flattening-shadow-dom\">Flattening Shadow DOM</a>.</p>\n<p>9.⠀<strong><code>screenshot</code></strong>, <strong><code>pdf</code></strong>, &amp; <strong><code>capture_mhtml</code></strong>:\n   - If <code>True</code>, captures a screenshot, PDF, or MHTML snapshot after the page is fully loaded.\n   - The results go to <code>result.screenshot</code> (base64), <code>result.pdf</code> (bytes), or <code>result.mhtml</code> (string).\n   - Use <code>force_viewport_screenshot=True</code> to capture only the visible viewport instead of the full page. This is faster and produces smaller images when you don't need a full-page screenshot.</p>\n<p>9.⠀<strong>Location Parameters</strong>:<br>\n   - <strong><code>locale</code></strong>: Browser's locale (e.g., <code>\"en-US\"</code>, <code>\"fr-FR\"</code>) for language preferences\n   - <strong><code>timezone_id</code></strong>: Browser's timezone (e.g., <code>\"America/New_York\"</code>, <code>\"Europe/Paris\"</code>)\n   - <strong><code>geolocation</code></strong>: GPS coordinates via <code>GeolocationConfig(latitude=48.8566, longitude=2.3522)</code>\n   - See <a href=\"../../advanced/identity-based-crawling/#7-locale-timezone-and-geolocation-control\">Identity Based Crawling</a></p>\n<p>10.⠀<strong>Proxy Configuration</strong>:\n    - <strong><code>proxy_config</code></strong>: Single <code>ProxyConfig</code> or <code>list[ProxyConfig]</code> — proxies tried in order. Pass a list for automatic escalation.\n    - <strong><code>proxy_rotation_strategy</code></strong>: Strategy for rotating proxies during crawls</p>\n<p>11.⠀<strong>Anti-Bot Retry &amp; Fallback</strong> (see <a href=\"../../advanced/anti-bot-and-fallback/\">Anti-Bot &amp; Fallback</a>):\n    - <strong><code>max_retries</code></strong>: Number of retry rounds when blocking is detected (default: 0). Each round tries all proxies in <code>proxy_config</code>.\n    - <strong><code>fallback_fetch_function</code></strong>: Async function called as last resort — takes URL, returns raw HTML</p>\n<p>12.⠀<strong>Page Interaction Parameters</strong>:\n    - <strong><code>scan_full_page</code></strong>: If <code>True</code>, scroll through the entire page to load all content\n    - <strong><code>wait_until</code></strong>: Condition to wait for when navigating (e.g., \"domcontentloaded\", \"networkidle\")\n    - <strong><code>page_timeout</code></strong>: Timeout in milliseconds for page operations (default: 60000)\n    - <strong><code>delay_before_return_html</code></strong>: Delay in seconds before retrieving final HTML.</p>\n<p>13.⠀<strong><code>url_matcher</code></strong> &amp; <strong><code>match_mode</code></strong>:<br>\n    - Enable URL-specific configurations when used with <code>arun_many()</code>.\n    - Set <code>url_matcher</code> to patterns (glob, function, or list) to match specific URLs.\n    - Use <code>match_mode</code> (OR/AND) to control how multiple patterns combine.\n    - See <a href=\"../../api/arun_many/#url-specific-configurations\">URL-Specific Configurations</a> for examples.</p>\n<p>13.⠀<strong><code>verbose</code></strong>:<br>\n    - Logs additional runtime details.<br>\n    - Overlaps with the browser's verbosity if also set to <code>True</code> in <code>BrowserConfig</code>.</p>\n<p>14.⠀<strong><code>stream</code></strong>:<br>\n    - If <code>True</code>, enables streaming mode for <code>arun_many()</code> to process URLs as they complete.\n    - Allows handling results incrementally instead of waiting for all URLs to finish.</p>\n<h3 id=\"helper-methods_1\">Helper Methods</h3>\n<p>The <code>clone()</code> method is particularly useful for creating variations of your crawler configuration:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Create a base configuration</span>\nbase_config = CrawlerRunConfig(\n    cache_mode=CacheMode.ENABLED,\n    word_count_threshold=200,\n    wait_until=<span class=\"hljs-string\">\"networkidle\"</span>\n)\n\n<span class=\"hljs-comment\"># Create variations for different use cases</span>\nstream_config = base_config.clone(\n    stream=True,  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    cache_mode=CacheMode.BYPASS\n)\n\ndebug_config = base_config.clone(\n    page_timeout=120000,  <span class=\"hljs-comment\"># Longer timeout for debugging</span>\n    verbose=True\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The <code>clone()</code> method:\n- Creates a new instance with all the same settings\n- Updates only the specified parameters\n- Leaves the original configuration unchanged\n- Perfect for creating variations without repeating all parameters</p>\n<hr>\n<h2 id=\"3-llmconfig-essentials\">3. LLMConfig Essentials</h2>\n<h3 id=\"key-fields-to-note_2\">Key fields to note</h3>\n<p>1.⠀<strong><code>provider</code></strong>:<br>\n- Which LLM provider to use. \n- Possible values are <code>\"ollama/llama3\",\"groq/llama3-70b-8192\",\"groq/llama3-8b-8192\", \"openai/gpt-4o-mini\" ,\"openai/gpt-4o\",\"openai/o1-mini\",\"openai/o1-preview\",\"openai/o3-mini\",\"openai/o3-mini-high\",\"anthropic/claude-3-haiku-20240307\",\"anthropic/claude-3-opus-20240229\",\"anthropic/claude-3-sonnet-20240229\",\"anthropic/claude-3-5-sonnet-20240620\",\"gemini/gemini-pro\",\"gemini/gemini-1.5-pro\",\"gemini/gemini-2.0-flash\",\"gemini/gemini-2.0-flash-exp\",\"gemini/gemini-2.0-flash-lite-preview-02-05\",\"deepseek/deepseek-chat\"</code><br><em>(default: <code>\"openai/gpt-4o-mini\"</code>)</em></p>\n<p>2.⠀<strong><code>api_token</code></strong>:<br>\n    - Optional. When not provided explicitly, api_token will be read from environment variables based on provider. For example: If a gemini model is passed as provider then,<code>\"GEMINI_API_KEY\"</code> will be read from environment variables<br>\n    - API token of LLM provider <br> eg: <code>api_token = \"gsk_1ClHGGJ7Lpn4WGybR7vNWGdyb3FY7zXEw3SCiy0BAVM9lL8CQv\"</code>\n    - Environment variable - use with prefix \"env:\" <br> eg:<code>api_token = \"env: GROQ_API_KEY\"</code>            </p>\n<p>3.⠀<strong><code>base_url</code></strong>:<br>\n   - If your provider has a custom endpoint</p>\n<p>4.⠀<strong>Retry/backoff controls</strong> <em>(optional)</em>:<br>\n   - <code>backoff_base_delay</code> <em>(default <code>2</code> seconds)</em> – base delay inserted before the first retry when the provider returns a rate-limit response.<br>\n   - <code>backoff_max_attempts</code> <em>(default <code>3</code>)</em> – total number of attempts (initial call plus retries) before the request is surfaced as an error.<br>\n   - <code>backoff_exponential_factor</code> <em>(default <code>2</code>)</em> – growth rate for the retry delay (<code>delay = base_delay * factor^attempt</code>).<br>\n   - These values are forwarded to the shared <code>perform_completion_with_backoff</code> helper, ensuring every strategy that consumes your <code>LLMConfig</code> honors the same throttling policy.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">llm_config = LLMConfig(\n    provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>,\n    api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>),\n    backoff_base_delay=1, <span class=\"hljs-comment\"># optional</span>\n    backoff_max_attempts=5, <span class=\"hljs-comment\"># optional</span>\n    backoff_exponential_factor=3, <span class=\"hljs-comment\">#optional</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"4-putting-it-all-together\">4. Putting It All Together</h2>\n<p>In a typical scenario, you define <strong>one</strong> <code>BrowserConfig</code> for your crawler session, then create <strong>one or more</strong> <code>CrawlerRunConfig</code> &amp; <code>LLMConfig</code> depending on each call's needs:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig, LLMContentFilter, DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1) Browser config: headless, bigger viewport, no proxy</span>\n    browser_conf = BrowserConfig(\n        headless=<span class=\"hljs-literal\">True</span>,\n        viewport_width=<span class=\"hljs-number\">1280</span>,\n        viewport_height=<span class=\"hljs-number\">720</span>\n    )\n\n    <span class=\"hljs-comment\"># 2) Example extraction strategy</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Articles\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.article\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n    extraction = JsonCssExtractionStrategy(schema)\n\n    <span class=\"hljs-comment\"># 3) Example LLM content filtering</span>\n\n    gemini_config = LLMConfig(\n        provider=<span class=\"hljs-string\">\"gemini/gemini-1.5-pro\"</span>, \n        api_token = <span class=\"hljs-string\">\"env:GEMINI_API_TOKEN\"</span>\n    )\n\n    <span class=\"hljs-comment\"># Initialize LLM filter with specific instruction</span>\n    <span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n        llm_config=gemini_config,  <span class=\"hljs-comment\"># or your preferred provider</span>\n        instruction=<span class=\"hljs-string\">\"\"\"\n        Focus on extracting the core educational content.\n        Include:\n        - Key concepts and explanations\n        - Important code examples\n        - Essential technical details\n        Exclude:\n        - Navigation elements\n        - Sidebars\n        - Footer content\n        Format the output as clean markdown with proper code blocks and headers.\n        \"\"\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">500</span>,  <span class=\"hljs-comment\"># Adjust based on your needs</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    md_generator = DefaultMarkdownGenerator(\n        content_filter=<span class=\"hljs-built_in\">filter</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># 4) Crawler run config: skip cache, use extraction</span>\n    run_conf = CrawlerRunConfig(\n        markdown_generator=md_generator,\n        extraction_strategy=extraction,\n        cache_mode=CacheMode.BYPASS,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4) Execute the crawl</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com/news\"</span>, config=run_conf)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted content:\"</span>, result.extracted_content)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"5-next-steps\">5. Next Steps</h2>\n<p>For a <strong>detailed list</strong> of available parameters (including advanced ones), see:</p>\n<ul>\n<li><a href=\"../../api/parameters/\">BrowserConfig, CrawlerRunConfig &amp; LLMConfig Reference</a>  </li>\n</ul>\n<p>You can explore topics like:</p>\n<ul>\n<li><strong>Custom Hooks &amp; Auth</strong> (Inject JavaScript or handle login forms).  </li>\n<li><strong>Session Management</strong> (Re-use pages, preserve state across multiple calls).  </li>\n<li><strong>Magic Mode</strong> or <strong>Identity-based Crawling</strong> (Fight bot detection by simulating user behavior).  </li>\n<li><strong>Advanced Caching</strong> (Fine-tune read/write cache modes).  </li>\n</ul>\n<hr>\n<h2 id=\"6-conclusion\">6. Conclusion</h2>\n<p><strong>BrowserConfig</strong>, <strong>CrawlerRunConfig</strong> and <strong>LLMConfig</strong> give you straightforward ways to define:</p>\n<ul>\n<li><strong>Which</strong> browser to launch, how it should run, and any proxy or user agent needs.  </li>\n<li><strong>How</strong> each crawl should behave—caching, timeouts, JavaScript code, extraction strategies, etc.</li>\n<li><strong>Which</strong> LLM provider to use, api token, temperature and base url for custom endpoints</li>\n</ul>\n<p>Use them together for <strong>clear, maintainable</strong> code, and when you need more specialized behavior, check out the advanced parameters in the <a href=\"../../api/parameters/\">reference docs</a>. Happy crawling!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/c4a-script/",
    "title": "C4A-Script - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"c4a-script-visual-web-automation-made-simple\">C4A-Script: Visual Web Automation Made Simple</h1>\n<h2 id=\"what-is-c4a-script\">What is C4A-Script?</h2>\n<p>C4A-Script is a powerful, human-readable domain-specific language (DSL) designed for web automation and interaction. Think of it as a simplified programming language that anyone can read and write, perfect for automating repetitive web tasks, testing user interfaces, or creating interactive demos.</p>\n<h3 id=\"why-c4a-script\">Why C4A-Script?</h3>\n<p><strong>Simple Syntax, Powerful Results</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Navigate and interact in plain English</span>\nGO https://example.com\nWAIT `<span class=\"hljs-comment\">#search-box` 5</span>\nTYPE <span class=\"hljs-string\">\"Hello World\"</span>\nCLICK `button[<span class=\"hljs-built_in\">type</span>=<span class=\"hljs-string\">\"submit\"</span>]`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Visual Programming Support</strong>\nC4A-Script comes with a built-in Blockly visual editor, allowing you to create scripts by dragging and dropping blocks - no coding experience required!</p>\n<p><strong>Perfect for:</strong>\n- <strong>UI Testing</strong>: Automate user interaction flows\n- <strong>Demo Creation</strong>: Build interactive product demonstrations<br>\n- <strong>Data Entry</strong>: Automate form filling and submissions\n- <strong>Testing Workflows</strong>: Validate complex user journeys\n- <strong>Training</strong>: Teach web automation without code complexity</p>\n<h2 id=\"getting-started-your-first-script\">Getting Started: Your First Script</h2>\n<p>Let's create a simple script that searches for something on a website:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># My first C4A-Script</span>\nGO <span class=\"hljs-symbol\">https</span><span class=\"hljs-punctuation\">:</span>//duckduckgo.com\n\n<span class=\"hljs-comment\"># Wait for the search box to appear</span>\nWAIT `<span class=\"hljs-keyword\">input</span><span class=\"hljs-punctuation\">[</span>name<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"q\"</span><span class=\"hljs-punctuation\">]</span>` <span class=\"hljs-number\">10</span>\n\n<span class=\"hljs-comment\"># Type our search query</span>\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"Crawl4AI\"</span>\n\n<span class=\"hljs-comment\"># Press Enter to search</span>\nPRESS Enter\n\n<span class=\"hljs-comment\"># Wait for results</span>\nWAIT `.results` <span class=\"hljs-number\">5</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>That's it! In just a few lines, you've automated a complete search workflow.</p>\n<h2 id=\"interactive-tutorial-live-demo\">Interactive Tutorial &amp; Live Demo</h2>\n<p>Want to learn by doing? We've got you covered:</p>\n<p><strong>🚀 <a href=\"https://docs.crawl4ai.com/apps/c4a-script/\">Live Demo</a></strong> - Try C4A-Script in your browser right now!</p>\n<p><strong>📁 <a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/c4a_script/\">Tutorial Examples</a></strong> - Complete examples with source code</p>\n<h3 id=\"running-the-tutorial-locally\">Running the Tutorial Locally</h3>\n<p>The tutorial includes a Flask-based web interface with:\n- <strong>Live Code Editor</strong> with syntax highlighting\n- <strong>Visual Blockly Editor</strong> for drag-and-drop programming\n- <strong>Recording Mode</strong> to capture your actions and generate scripts\n- <strong>Timeline View</strong> to see and edit your automation steps</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Clone and navigate to the tutorial</span>\n<span class=\"hljs-built_in\">cd</span> docs/examples/c4a_script/tutorial/\n\n<span class=\"hljs-comment\"># Install dependencies</span>\npip install -r requirements.txt\n\n<span class=\"hljs-comment\"># Launch the tutorial server</span>\npython server.py\n\n<span class=\"hljs-comment\"># Open http://localhost:8000 in your browser</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"core-concepts\">Core Concepts</h2>\n<h3 id=\"commands-and-syntax\">Commands and Syntax</h3>\n<p>C4A-Script uses simple, English-like commands. Each command does one specific thing:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Comments start with #</span>\nCOMMAND parameter1 parameter2\n\n<span class=\"hljs-comment\"># Most commands use CSS selectors in backticks</span>\nCLICK `<span class=\"hljs-comment\">#submit-button`</span>\n\n<span class=\"hljs-comment\"># Text content goes in quotes</span>\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"Hello, World!\"</span>\n\n<span class=\"hljs-comment\"># Numbers are used directly</span>\nWAIT <span class=\"hljs-number\">3</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"selectors-finding-elements\">Selectors: Finding Elements</h3>\n<p>C4A-Script uses CSS selectors to identify elements on the page:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># By ID</span>\nCLICK `<span class=\"hljs-comment\">#login-button`</span>\n\n<span class=\"hljs-comment\"># By class</span>\nCLICK `.submit-btn`\n\n<span class=\"hljs-comment\"># By attribute</span>\nCLICK `button<span class=\"hljs-punctuation\">[</span><span class=\"hljs-keyword\">type</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"submit\"</span><span class=\"hljs-punctuation\">]</span>`\n\n<span class=\"hljs-comment\"># By accessible attributes</span>\nCLICK `button<span class=\"hljs-punctuation\">[</span>aria-label<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Search\"</span><span class=\"hljs-punctuation\">]</span><span class=\"hljs-punctuation\">[</span>title<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Search\"</span><span class=\"hljs-punctuation\">]</span>`\n\n<span class=\"hljs-comment\"># Complex selectors</span>\nCLICK `.form-container <span class=\"hljs-keyword\">input</span><span class=\"hljs-punctuation\">[</span>name<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"email\"</span><span class=\"hljs-punctuation\">]</span>`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"variables-and-dynamic-content\">Variables and Dynamic Content</h3>\n<p>Store and reuse values with variables:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Set a variable</span>\nSETVAR username = <span class=\"hljs-string\">\"john@example.com\"</span>\nSETVAR password = <span class=\"hljs-string\">\"secret123\"</span>\n\n<span class=\"hljs-comment\"># Use variables (prefix with $)</span>\nTYPE $username\nPRESS Tab\nTYPE $password\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"command-categories\">Command Categories</h2>\n<h3 id=\"navigation-commands\">🧭 Navigation Commands</h3>\n<p>Move around the web like a user would:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>GO</code></td>\n<td>Navigate to URL</td>\n<td><code>GO https://example.com</code></td>\n</tr>\n<tr>\n<td><code>RELOAD</code></td>\n<td>Refresh current page</td>\n<td><code>RELOAD</code></td>\n</tr>\n<tr>\n<td><code>BACK</code></td>\n<td>Go back in history</td>\n<td><code>BACK</code></td>\n</tr>\n<tr>\n<td><code>FORWARD</code></td>\n<td>Go forward in history</td>\n<td><code>FORWARD</code></td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"wait-commands\">⏱️ Wait Commands</h3>\n<p>Ensure elements are ready before interacting:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>WAIT</code></td>\n<td>Wait for time/element/text</td>\n<td><code>WAIT 3</code> or <code>WAIT \\</code>#element` 10`</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"mouse-commands\">🖱️ Mouse Commands</h3>\n<p>Click, drag, and move like a human:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>CLICK</code></td>\n<td>Click element or coordinates</td>\n<td><code>CLICK \\</code>button`<code>or</code>CLICK 100 200`</td>\n</tr>\n<tr>\n<td><code>DOUBLE_CLICK</code></td>\n<td>Double-click element</td>\n<td><code>DOUBLE_CLICK \\</code>.item``</td>\n</tr>\n<tr>\n<td><code>RIGHT_CLICK</code></td>\n<td>Right-click element</td>\n<td><code>RIGHT_CLICK \\</code>#menu``</td>\n</tr>\n<tr>\n<td><code>SCROLL</code></td>\n<td>Scroll in direction</td>\n<td><code>SCROLL DOWN 500</code></td>\n</tr>\n<tr>\n<td><code>DRAG</code></td>\n<td>Drag from point to point</td>\n<td><code>DRAG 100 100 500 300</code></td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"keyboard-commands\">⌨️ Keyboard Commands</h3>\n<p>Type text and press keys naturally:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>TYPE</code></td>\n<td>Type text or variable</td>\n<td><code>TYPE \"Hello\"</code> or <code>TYPE $username</code></td>\n</tr>\n<tr>\n<td><code>PRESS</code></td>\n<td>Press special keys</td>\n<td><code>PRESS Tab</code> or <code>PRESS Enter</code></td>\n</tr>\n<tr>\n<td><code>CLEAR</code></td>\n<td>Clear input field</td>\n<td><code>CLEAR \\</code>#search``</td>\n</tr>\n<tr>\n<td><code>SET</code></td>\n<td>Set input value directly</td>\n<td><code>SET \\</code>#email` \"user@example.com\"`</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"control-flow\">🔀 Control Flow</h3>\n<p>Add logic and repetition to your scripts:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>IF</code></td>\n<td>Conditional execution</td>\n<td><code>IF (EXISTS \\</code>#popup`) THEN CLICK `#close``</td>\n</tr>\n<tr>\n<td><code>REPEAT</code></td>\n<td>Loop commands</td>\n<td><code>REPEAT (SCROLL DOWN 300, 5)</code></td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"variables-advanced\">💾 Variables &amp; Advanced</h3>\n<p>Store data and execute custom code:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Command</th>\n<th>Purpose</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>SETVAR</code></td>\n<td>Create variable</td>\n<td><code>SETVAR email = \"test@example.com\"</code></td>\n</tr>\n<tr>\n<td><code>EVAL</code></td>\n<td>Execute JavaScript</td>\n<td><code>EVAL \\</code>console.log('Hello')``</td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"real-world-examples\">Real-World Examples</h2>\n<h3 id=\"example-1-login-flow\">Example 1: Login Flow</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Complete login automation</span>\nGO <span class=\"hljs-symbol\">https</span><span class=\"hljs-punctuation\">:</span>//myapp.com/login\n\n<span class=\"hljs-comment\"># Wait for page to load</span>\nWAIT `<span class=\"hljs-comment\">#login-form` 5</span>\n\n<span class=\"hljs-comment\"># Fill credentials</span>\nCLICK `<span class=\"hljs-comment\">#email`</span>\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"user@example.com\"</span>\nPRESS Tab\n<span class=\"hljs-keyword\">TYPE</span> <span class=\"hljs-string\">\"mypassword\"</span>\n\n<span class=\"hljs-comment\"># Submit form</span>\nCLICK `button<span class=\"hljs-punctuation\">[</span><span class=\"hljs-keyword\">type</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"submit\"</span><span class=\"hljs-punctuation\">]</span>`\n\n<span class=\"hljs-comment\"># Wait for dashboard</span>\nWAIT `.dashboard` <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"example-2-e-commerce-shopping\">Example 2: E-commerce Shopping</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Shopping automation with variables</span>\nSETVAR product = <span class=\"hljs-string\">\"laptop\"</span>\nSETVAR budget = <span class=\"hljs-string\">\"1000\"</span>\n\nGO https:<span class=\"hljs-regexp\">//s</span>hop.example.com\nWAIT <span class=\"hljs-string\">`#search-box`</span> <span class=\"hljs-number\">3</span>\n\n<span class=\"hljs-comment\"># Search for product</span>\nTYPE $product\nPRESS Enter\nWAIT <span class=\"hljs-string\">`.product-list`</span> <span class=\"hljs-number\">5</span>\n\n<span class=\"hljs-comment\"># Filter by price</span>\nCLICK <span class=\"hljs-string\">`.price-filter`</span>\nSET <span class=\"hljs-string\">`#max-price`</span> $budget\nCLICK <span class=\"hljs-string\">`.apply-filters`</span>\n\n<span class=\"hljs-comment\"># Select first result</span>\nWAIT <span class=\"hljs-string\">`.product-item`</span> <span class=\"hljs-number\">3</span>\nCLICK <span class=\"hljs-string\">`.product-item:first-child`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"example-3-form-automation-with-conditions\">Example 3: Form Automation with Conditions</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Smart form filling with error handling</span>\nGO https:<span class=\"hljs-regexp\">//</span>forms.example.com\n\n<span class=\"hljs-comment\"># Check if user is already logged in</span>\nIF (EXISTS <span class=\"hljs-string\">`.user-menu`</span>) THEN GO https:<span class=\"hljs-regexp\">//</span>forms.example.com/new\nIF (NOT EXISTS <span class=\"hljs-string\">`.user-menu`</span>) THEN CLICK <span class=\"hljs-string\">`#login-link`</span>\n\n<span class=\"hljs-comment\"># Fill form</span>\nWAIT <span class=\"hljs-string\">`#contact-form`</span> <span class=\"hljs-number\">5</span>\nSET <span class=\"hljs-string\">`#name`</span> <span class=\"hljs-string\">\"John Doe\"</span>\nSET <span class=\"hljs-string\">`#email`</span> <span class=\"hljs-string\">\"john@example.com\"</span>\nSET <span class=\"hljs-string\">`#message`</span> <span class=\"hljs-string\">\"Hello from C4A-Script!\"</span>\n\n<span class=\"hljs-comment\"># Handle popup if it appears</span>\nIF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept-cookies`</span>\n\n<span class=\"hljs-comment\"># Submit</span>\nCLICK <span class=\"hljs-string\">`#submit-button`</span>\nWAIT <span class=\"hljs-string\">`.success-message`</span> <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"visual-programming-with-blockly\">Visual Programming with Blockly</h2>\n<p>C4A-Script includes a powerful visual programming interface built on Google Blockly. Perfect for:</p>\n<ul>\n<li><strong>Non-programmers</strong> who want to create automation</li>\n<li><strong>Rapid prototyping</strong> of automation workflows  </li>\n<li><strong>Educational environments</strong> for teaching automation concepts</li>\n<li><strong>Collaborative development</strong> where visual representation helps communication</li>\n</ul>\n<h3 id=\"features\">Features:</h3>\n<ul>\n<li><strong>Drag &amp; Drop Interface</strong>: Build scripts by connecting blocks</li>\n<li><strong>Real-time Sync</strong>: Changes in visual mode instantly update the text script</li>\n<li><strong>Smart Block Types</strong>: Blocks are categorized by function (Navigation, Actions, etc.)</li>\n<li><strong>Error Prevention</strong>: Visual connections prevent syntax errors</li>\n<li><strong>Comment Support</strong>: Add visual comment blocks for documentation</li>\n</ul>\n<p>Try the visual editor in our <a href=\"https://docs.crawl4ai.com/c4a-script/demo\">live demo</a> or <a href=\"/examples/c4a_script/tutorial/\">local tutorial</a>.</p>\n<h2 id=\"advanced-features\">Advanced Features</h2>\n<h3 id=\"recording-mode\">Recording Mode</h3>\n<p>The tutorial interface includes a recording feature that watches your browser interactions and automatically generates C4A-Script commands:</p>\n<ol>\n<li>Click \"Record\" in the tutorial interface</li>\n<li>Perform actions in the browser preview</li>\n<li>Watch as C4A-Script commands are generated in real-time</li>\n<li>Edit and refine the generated script</li>\n</ol>\n<h3 id=\"error-handling-and-debugging\">Error Handling and Debugging</h3>\n<p>C4A-Script provides clear error messages and debugging information:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Use comments for debugging</span>\n<span class=\"hljs-comment\"># This will wait up to 10 seconds for the element</span>\nWAIT `<span class=\"hljs-comment\">#slow-loading-element` 10</span>\n\n<span class=\"hljs-comment\"># Check if element exists before clicking</span>\nIF <span class=\"hljs-punctuation\">(</span>EXISTS `<span class=\"hljs-comment\">#optional-button`) THEN CLICK `#optional-button`</span>\n\n<span class=\"hljs-comment\"># Use EVAL for custom debugging</span>\nEVAL `console.log<span class=\"hljs-punctuation\">(</span><span class=\"hljs-string\">\"Current page title:\"</span>, document.title<span class=\"hljs-punctuation\">)</span>`\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"integration-with-crawl4ai\">Integration with Crawl4AI</h3>\n<p>C4A-Script integrates seamlessly with Crawl4AI's web crawling capabilities:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Use C4A-Script for interaction before crawling</span>\nscript = <span class=\"hljs-string\">\"\"\"\nGO https://example.com\nCLICK `#load-more-content`\nWAIT `.dynamic-content` 5\n\"\"\"</span>\n\nconfig = CrawlerRunConfig(\n    js_code=script,\n    wait_for=<span class=\"hljs-string\">\".dynamic-content\"</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n    <span class=\"hljs-built_in\">print</span>(result.markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices\">Best Practices</h2>\n<h3 id=\"1-always-wait-for-elements\">1. Always Wait for Elements</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Bad: Clicking immediately</span>\nCLICK `<span class=\"hljs-comment\">#button`</span>\n\n<span class=\"hljs-comment\"># Good: Wait for element to appear</span>\nWAIT `<span class=\"hljs-comment\">#button` 5</span>\nCLICK `<span class=\"hljs-comment\">#button`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-use-descriptive-comments\">2. Use Descriptive Comments</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Login to user account</span>\nGO https:<span class=\"hljs-regexp\">//m</span>yapp.com/login\nWAIT <span class=\"hljs-string\">`#login-form`</span> <span class=\"hljs-number\">5</span>\n\n<span class=\"hljs-comment\"># Enter credentials</span>\nTYPE <span class=\"hljs-string\">\"user@example.com\"</span>\nPRESS Tab\nTYPE <span class=\"hljs-string\">\"password123\"</span>\n\n<span class=\"hljs-comment\"># Submit and wait for redirect</span>\nCLICK <span class=\"hljs-string\">`#submit-button`</span>\nWAIT <span class=\"hljs-string\">`.dashboard`</span> <span class=\"hljs-number\">10</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-handle-variable-conditions\">3. Handle Variable Conditions</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Handle different page states</span>\nIF (EXISTS <span class=\"hljs-string\">`.cookie-banner`</span>) THEN CLICK <span class=\"hljs-string\">`.accept-cookies`</span>\nIF (EXISTS <span class=\"hljs-string\">`.popup-modal`</span>) THEN CLICK <span class=\"hljs-string\">`.close-modal`</span>\n\n<span class=\"hljs-comment\"># Proceed with main workflow</span>\nCLICK <span class=\"hljs-string\">`#main-action`</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"4-use-variables-for-reusability\">4. Use Variables for Reusability</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Define once, use everywhere</span>\nSETVAR base_url = <span class=\"hljs-string\">\"https://myapp.com\"</span>\nSETVAR test_email = <span class=\"hljs-string\">\"test@example.com\"</span>\n\nGO $base_url/login\nSET <span class=\"hljs-string\">`#email`</span> $test_email\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"getting-help\">Getting Help</h2>\n<ul>\n<li><strong>📖 <a href=\"/examples/c4a_script/\">Complete Examples</a></strong> - Real-world automation scripts</li>\n<li><strong>🎮 <a href=\"/examples/c4a_script/tutorial/\">Interactive Tutorial</a></strong> - Hands-on learning environment  </li>\n<li><strong>📋 <a href=\"/api/c4a-script-reference/\">API Reference</a></strong> - Detailed command documentation</li>\n<li><strong>🌐 <a href=\"https://docs.crawl4ai.com/c4a-script/demo\">Live Demo</a></strong> - Try it in your browser</li>\n</ul>\n<h2 id=\"whats-next\">What's Next?</h2>\n<p>Ready to dive deeper? Check out:</p>\n<ol>\n<li><strong><a href=\"/api/c4a-script-reference/\">API Reference</a></strong> - Complete command documentation</li>\n<li><strong><a href=\"/examples/c4a_script/\">Tutorial Examples</a></strong> - Copy-paste ready scripts</li>\n<li><strong><a href=\"/examples/c4a_script/tutorial/\">Local Tutorial Setup</a></strong> - Run the full development environment</li>\n</ol>\n<p>C4A-Script makes web automation accessible to everyone. Whether you're a developer automating tests, a designer creating interactive demos, or a business user streamlining repetitive tasks, C4A-Script has the tools you need.</p>\n<p><em>Start automating today - your future self will thank you!</em> 🚀</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/cache-modes/",
    "title": "Cache Modes - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-cache-system-and-migration-guide\">Crawl4AI Cache System and Migration Guide</h1>\n<h2 id=\"overview\">Overview</h2>\n<p>Starting from version 0.5.0, Crawl4AI introduces a new caching system that replaces the old boolean flags with a more intuitive <code>CacheMode</code> enum. This change simplifies cache control and makes the behavior more predictable.</p>\n<h2 id=\"old-vs-new-approach\">Old vs New Approach</h2>\n<h3 id=\"old-way-deprecated\">Old Way (Deprecated)</h3>\n<p>The old system used multiple boolean flags:\n- <code>bypass_cache</code>: Skip cache entirely\n- <code>disable_cache</code>: Disable all caching\n- <code>no_cache_read</code>: Don't read from cache\n- <code>no_cache_write</code>: Don't write to cache</p>\n<h3 id=\"new-way-recommended\">New Way (Recommended)</h3>\n<p>The new system uses a single <code>CacheMode</code> enum:\n- <code>CacheMode.ENABLED</code>: Normal caching (read/write)\n- <code>CacheMode.DISABLED</code>: No caching at all\n- <code>CacheMode.READ_ONLY</code>: Only read from cache\n- <code>CacheMode.WRITE_ONLY</code>: Only write to cache\n- <code>CacheMode.BYPASS</code>: Skip cache for this operation</p>\n<h2 id=\"migration-example\">Migration Example</h2>\n<h3 id=\"old-code-deprecated\">Old Code (Deprecated)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">old_code</span>(<span class=\"hljs-params\">crawler: AsyncWebCrawler</span>):\n    <span class=\"hljs-comment\"># Legacy `bypass_cache` / `disable_cache` / `no_cache_read` / `no_cache_write`</span>\n    <span class=\"hljs-comment\"># were removed in v0.5+. This example no longer applies:</span>\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://www.nbcnews.com/business\"</span>,\n        <span class=\"hljs-comment\"># cache_mode is the only cache option now.</span>\n    )\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.markdown))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"new-code-recommended\">New Code (Recommended)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">use_proxy</span>():\n    <span class=\"hljs-comment\"># Use CacheMode in CrawlerRunConfig</span>\n    config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)  \n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.nbcnews.com/business\"</span>,\n            config=config  <span class=\"hljs-comment\"># Pass the configuration object</span>\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.markdown))\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">await</span> use_proxy()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"common-migration-patterns\">Common Migration Patterns</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Legacy Flag</th>\n<th>Replacement</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>bypass_cache</code></td>\n<td><code>cache_mode=CacheMode.BYPASS</code></td>\n</tr>\n<tr>\n<td><code>disable_cache</code></td>\n<td><code>cache_mode=CacheMode.DISABLED</code></td>\n</tr>\n<tr>\n<td><code>no_cache_read</code></td>\n<td><code>cache_mode=CacheMode.READ_ONLY</code></td>\n</tr>\n<tr>\n<td><code>no_cache_write</code></td>\n<td><code>cache_mode=CacheMode.WRITE_ONLY</code></td>\n</tr>\n</tbody>\n</table>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/cli/",
    "title": "Command Line Interface - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl4ai-cli-guide\">Crawl4AI CLI Guide</h1>\n<h2 id=\"table-of-contents\">Table of Contents</h2>\n<ul>\n<li><a href=\"#installation\">Installation</a></li>\n<li><a href=\"#basic-usage\">Basic Usage</a></li>\n<li><a href=\"#configuration\">Configuration</a></li>\n<li><a href=\"#browser-configuration\">Browser Configuration</a></li>\n<li><a href=\"#crawler-configuration\">Crawler Configuration</a></li>\n<li><a href=\"#extraction-configuration\">Extraction Configuration</a></li>\n<li><a href=\"#content-filtering\">Content Filtering</a></li>\n<li><a href=\"#advanced-features\">Advanced Features</a></li>\n<li><a href=\"#llm-qa\">LLM Q&amp;A</a></li>\n<li><a href=\"#structured-data-extraction\">Structured Data Extraction</a></li>\n<li><a href=\"#content-filtering-1\">Content Filtering</a></li>\n<li><a href=\"#output-formats\">Output Formats</a></li>\n<li><a href=\"#examples\">Examples</a></li>\n<li><a href=\"#configuration-reference\">Configuration Reference</a></li>\n<li><a href=\"#best-practices--tips\">Best Practices &amp; Tips</a></li>\n</ul>\n<h2 id=\"installation\">Installation</h2>\n<p>The Crawl4AI CLI will be installed automatically when you install the library.</p>\n<h2 id=\"basic-usage\">Basic Usage</h2>\n<p>The Crawl4AI CLI (<code>crwl</code>) provides a simple interface to the Crawl4AI library:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Basic crawling</span>\ncrwl https://example.com\n\n<span class=\"hljs-comment\"># Get markdown output</span>\ncrwl https://example.com -o markdown\n\n<span class=\"hljs-comment\"># Verbose JSON output with cache bypass</span>\ncrwl https://example.com -o json -v --bypass-cache\n\n<span class=\"hljs-comment\"># See usage examples</span>\ncrwl --example\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"quick-example-of-advanced-usage\">Quick Example of Advanced Usage</h2>\n<p>If you clone the repository and run the following command, you will receive the content of the page in JSON format according to a JSON-CSS schema:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-swift\">crwl <span class=\"hljs-string\">\"https://www.infoq.com/ai-ml-data-eng/\"</span> <span class=\"hljs-operator\">-</span>e docs<span class=\"hljs-regexp\">/examples/</span>cli<span class=\"hljs-regexp\">/extract_css.yml -s docs/</span>examples<span class=\"hljs-regexp\">/cli/</span>css_schema.json <span class=\"hljs-operator\">-</span>o json;\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"configuration\">Configuration</h2>\n<h3 id=\"browser-configuration\">Browser Configuration</h3>\n<p>Browser settings can be configured via YAML file or command line parameters:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-yaml\"><span class=\"hljs-comment\"># browser.yml</span>\n<span class=\"hljs-attr\">headless:</span> <span class=\"hljs-literal\">true</span>\n<span class=\"hljs-attr\">viewport_width:</span> <span class=\"hljs-number\">1280</span>\n<span class=\"hljs-attr\">user_agent_mode:</span> <span class=\"hljs-string\">\"random\"</span>\n<span class=\"hljs-attr\">verbose:</span> <span class=\"hljs-literal\">true</span>\n<span class=\"hljs-attr\">ignore_https_errors:</span> <span class=\"hljs-literal\">true</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Using config file</span>\ncrwl https://example.com -B browser.yml\n\n<span class=\"hljs-comment\"># Using direct parameters</span>\ncrwl https://example.com -b <span class=\"hljs-string\">\"headless=true,viewport_width=1280,user_agent_mode=random\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"crawler-configuration\">Crawler Configuration</h3>\n<p>Control crawling behavior:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-vbnet\"># crawler.yml\n<span class=\"hljs-symbol\">cache_mode:</span> <span class=\"hljs-string\">\"bypass\"</span>\n<span class=\"hljs-symbol\">wait_until:</span> <span class=\"hljs-string\">\"networkidle\"</span>\n<span class=\"hljs-symbol\">page_timeout:</span> <span class=\"hljs-number\">30000</span>\n<span class=\"hljs-symbol\">delay_before_return_html:</span> <span class=\"hljs-number\">0.5</span>\n<span class=\"hljs-symbol\">word_count_threshold:</span> <span class=\"hljs-number\">100</span>\n<span class=\"hljs-symbol\">scan_full_page:</span> <span class=\"hljs-literal\">true</span>\n<span class=\"hljs-symbol\">scroll_delay:</span> <span class=\"hljs-number\">0.3</span>\n<span class=\"hljs-symbol\">process_iframes:</span> <span class=\"hljs-literal\">false</span>\n<span class=\"hljs-symbol\">remove_overlay_elements:</span> <span class=\"hljs-literal\">true</span>\n<span class=\"hljs-symbol\">magic:</span> <span class=\"hljs-literal\">true</span>\n<span class=\"hljs-symbol\">verbose:</span> <span class=\"hljs-literal\">true</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Using config file</span>\ncrwl https://example.com -C crawler.yml\n\n<span class=\"hljs-comment\"># Using direct parameters</span>\ncrwl https://example.com -c <span class=\"hljs-string\">\"css_selector=#main,delay_before_return_html=2,scan_full_page=true\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"extraction-configuration\">Extraction Configuration</h3>\n<p>Two types of extraction are supported:</p>\n<ol>\n<li>CSS/XPath-based extraction:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-yaml\"><span class=\"hljs-comment\"># extract_css.yml</span>\n<span class=\"hljs-attr\">type:</span> <span class=\"hljs-string\">\"json-css\"</span>\n<span class=\"hljs-attr\">params:</span>\n  <span class=\"hljs-attr\">verbose:</span> <span class=\"hljs-literal\">true</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ol>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-comment\">// css_schema.json</span>\n<span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"ArticleExtractor\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\".article\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-punctuation\">{</span>\n      <span class=\"hljs-attr\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h1.title\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n    <span class=\"hljs-punctuation\">}</span><span class=\"hljs-punctuation\">,</span>\n    <span class=\"hljs-punctuation\">{</span>\n      <span class=\"hljs-attr\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"link\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"a.read-more\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span>\n    <span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ol>\n<li>LLM-based extraction:\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-vbnet\"># extract_llm.yml\n<span class=\"hljs-symbol\">type:</span> <span class=\"hljs-string\">\"llm\"</span>\n<span class=\"hljs-symbol\">provider:</span> <span class=\"hljs-string\">\"openai/gpt-4\"</span>\n<span class=\"hljs-symbol\">instruction:</span> <span class=\"hljs-string\">\"Extract all articles with their titles and links\"</span>\n<span class=\"hljs-symbol\">api_token:</span> <span class=\"hljs-string\">\"your-token\"</span>\n<span class=\"hljs-symbol\">params:</span>\n  temperature: <span class=\"hljs-number\">0.3</span>\n  max_tokens: <span class=\"hljs-number\">1000</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div></li>\n</ol>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-comment\">// llm_schema.json</span>\n<span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"title\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Article\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"object\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"properties\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-attr\">\"title\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span>\n      <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"string\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"description\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"The title of the article\"</span>\n    <span class=\"hljs-punctuation\">}</span><span class=\"hljs-punctuation\">,</span>\n    <span class=\"hljs-attr\">\"link\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span>\n      <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"string\"</span><span class=\"hljs-punctuation\">,</span>\n      <span class=\"hljs-attr\">\"description\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"URL to the full article\"</span>\n    <span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">}</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"advanced-features\">Advanced Features</h2>\n<h3 id=\"llm-qa\">LLM Q&amp;A</h3>\n<p>Ask questions about crawled content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Simple question</span>\ncrwl https://example.com -q <span class=\"hljs-string\">\"What is the main topic discussed?\"</span>\n\n<span class=\"hljs-comment\"># View content then ask questions</span>\ncrwl https://example.com -o markdown  <span class=\"hljs-comment\"># See content first</span>\ncrwl https://example.com -q <span class=\"hljs-string\">\"Summarize the key points\"</span>\ncrwl https://example.com -q <span class=\"hljs-string\">\"What are the conclusions?\"</span>\n\n<span class=\"hljs-comment\"># Combined with advanced crawling</span>\ncrwl https://example.com \\\n    -B browser.yml \\\n    -c <span class=\"hljs-string\">\"css_selector=article,scan_full_page=true\"</span> \\\n    -q <span class=\"hljs-string\">\"What are the pros and cons mentioned?\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>First-time setup:\n- Prompts for LLM provider and API token\n- Saves configuration in <code>~/.crawl4ai/global.yml</code>\n- Supports various providers (openai/gpt-4, anthropic/claude-3-sonnet, etc.)\n- For case of <code>ollama</code> you do not need to provide API token.\n- See <a href=\"https://docs.litellm.ai/docs/providers\">LiteLLM Providers</a> for full list</p>\n<h3 id=\"structured-data-extraction\">Structured Data Extraction</h3>\n<p>Extract structured data using CSS selectors:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">crwl https:<span class=\"hljs-comment\">//example.com \\\n    -e extract_css.yml \\\n    -s css_schema.json \\\n    -o json</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Or using LLM-based extraction:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">crwl https:<span class=\"hljs-comment\">//example.com \\\n    -e extract_llm.yml \\\n    -s llm_schema.json \\\n    -o json</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"content-filtering\">Content Filtering</h3>\n<p>Filter content for relevance:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-vbnet\"># filter_bm25.yml\n<span class=\"hljs-symbol\">type:</span> <span class=\"hljs-string\">\"bm25\"</span>\n<span class=\"hljs-symbol\">query:</span> <span class=\"hljs-string\">\"target content\"</span>\n<span class=\"hljs-symbol\">threshold:</span> <span class=\"hljs-number\">1.0</span>\n\n# filter_pruning.yml\n<span class=\"hljs-symbol\">type:</span> <span class=\"hljs-string\">\"pruning\"</span>\n<span class=\"hljs-symbol\">query:</span> <span class=\"hljs-string\">\"focus topic\"</span>\n<span class=\"hljs-symbol\">threshold:</span> <span class=\"hljs-number\">0.48</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">crwl https://example.com -f filter_bm25.yml -o markdown-fit\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"output-formats\">Output Formats</h2>\n<ul>\n<li><code>all</code> - Full crawl result including metadata</li>\n<li><code>json</code> - Extracted structured data (when using extraction)</li>\n<li><code>markdown</code> / <code>md</code> - Raw markdown output</li>\n<li><code>markdown-fit</code> / <code>md-fit</code> - Filtered markdown for better readability</li>\n</ul>\n<h2 id=\"complete-examples\">Complete Examples</h2>\n<ol>\n<li>\n<p>Basic Extraction:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">crwl https:<span class=\"hljs-comment\">//example.com \\\n    -B browser.yml \\\n    -C crawler.yml \\\n    -o json</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p>Structured Data Extraction:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">crwl https:<span class=\"hljs-comment\">//example.com \\\n    -e extract_css.yml \\\n    -s css_schema.json \\\n    -o json \\\n    -v</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p>LLM Extraction with Filtering:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">crwl https:<span class=\"hljs-comment\">//example.com \\\n    -B browser.yml \\\n    -e extract_llm.yml \\\n    -s llm_schema.json \\\n    -f filter_bm25.yml \\\n    -o json</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p>Interactive Q&amp;A:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># First crawl and view</span>\ncrwl https://example.com -o markdown\n\n<span class=\"hljs-comment\"># Then ask questions</span>\ncrwl https://example.com -q <span class=\"hljs-string\">\"What are the main points?\"</span>\ncrwl https://example.com -q <span class=\"hljs-string\">\"Summarize the conclusions\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n<h2 id=\"best-practices-tips\">Best Practices &amp; Tips</h2>\n<ol>\n<li><strong>Configuration Management</strong>:</li>\n<li>Keep common configurations in YAML files</li>\n<li>Use CLI parameters for quick overrides</li>\n<li>\n<p>Store sensitive data (API tokens) in <code>~/.crawl4ai/global.yml</code></p>\n</li>\n<li>\n<p><strong>Performance Optimization</strong>:</p>\n</li>\n<li>Use <code>--bypass-cache</code> for fresh content</li>\n<li>Enable <code>scan_full_page</code> for infinite scroll pages</li>\n<li>\n<p>Adjust <code>delay_before_return_html</code> for dynamic content</p>\n</li>\n<li>\n<p><strong>Content Extraction</strong>:</p>\n</li>\n<li>Use CSS extraction for structured content</li>\n<li>Use LLM extraction for unstructured content</li>\n<li>\n<p>Combine with filters for focused results</p>\n</li>\n<li>\n<p><strong>Q&amp;A Workflow</strong>:</p>\n</li>\n<li>View content first with <code>-o markdown</code></li>\n<li>Ask specific questions</li>\n<li>Use broader context with appropriate selectors</li>\n</ol>\n<h2 id=\"recap\">Recap</h2>\n<p>The Crawl4AI CLI provides:\n- Flexible configuration via files and parameters\n- Multiple extraction strategies (CSS, XPath, LLM)\n- Content filtering and optimization\n- Interactive Q&amp;A capabilities\n- Various output formats</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/content-selection/",
    "title": "Content Selection - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"content-selection\">Content Selection</h1>\n<p>Crawl4AI provides multiple ways to <strong>select</strong>, <strong>filter</strong>, and <strong>refine</strong> the content from your crawls. Whether you need to target a specific CSS region, exclude entire tags, filter out external links, or remove certain domains and images, <strong><code>CrawlerRunConfig</code></strong> offers a wide range of parameters.</p>\n<p>Below, we show how to configure these parameters and combine them for precise control.</p>\n<hr>\n<h2 id=\"1-css-based-selection\">1. CSS-Based Selection</h2>\n<p>There are two ways to select content from a page: using <code>css_selector</code> or the more flexible <code>target_elements</code>.</p>\n<h3 id=\"11-using-css_selector\">1.1 Using <code>css_selector</code></h3>\n<p>A straightforward way to <strong>limit</strong> your crawl results to a certain region of the page is <strong><code>css_selector</code></strong> in <strong><code>CrawlerRunConfig</code></strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># e.g., first 30 items from Hacker News</span>\n        css_selector=<span class=\"hljs-string\">\".athing:nth-child(-n+30)\"</span>  \n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Partial HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Result</strong>: Only elements matching that selector remain in <code>result.cleaned_html</code>.</p>\n<h3 id=\"12-using-target_elements\">1.2 Using <code>target_elements</code></h3>\n<p>The <code>target_elements</code> parameter provides more flexibility by allowing you to target <strong>multiple elements</strong> for content extraction while preserving the entire page context for other features:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Target article body and sidebar, but not other content</span>\n        target_elements=[<span class=\"hljs-string\">\"article.main-content\"</span>, <span class=\"hljs-string\">\"aside.sidebar\"</span>]\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/blog-post\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown focused on target elements\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Links from entire page still available:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key difference</strong>: With <code>target_elements</code>, the markdown generation and structural data extraction focus on those elements, but other page elements (like links, images, and tables) are still extracted from the entire page. This gives you fine-grained control over what appears in your markdown content while preserving full page context for link analysis and media collection.</p>\n<hr>\n<h2 id=\"2-content-filtering-exclusions\">2. Content Filtering &amp; Exclusions</h2>\n<h3 id=\"21-basic-overview\">2.1 Basic Overview</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    <span class=\"hljs-comment\"># Content thresholds</span>\n    word_count_threshold=<span class=\"hljs-number\">10</span>,        <span class=\"hljs-comment\"># Minimum words per block</span>\n\n    <span class=\"hljs-comment\"># Tag exclusions</span>\n    excluded_tags=[<span class=\"hljs-string\">'form'</span>, <span class=\"hljs-string\">'header'</span>, <span class=\"hljs-string\">'footer'</span>, <span class=\"hljs-string\">'nav'</span>],\n\n    <span class=\"hljs-comment\"># Link filtering</span>\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,    \n    exclude_social_media_links=<span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-comment\"># Block entire domains</span>\n    exclude_domains=[<span class=\"hljs-string\">\"adtrackers.com\"</span>, <span class=\"hljs-string\">\"spammynews.org\"</span>],    \n    exclude_social_media_domains=[<span class=\"hljs-string\">\"facebook.com\"</span>, <span class=\"hljs-string\">\"twitter.com\"</span>],\n\n    <span class=\"hljs-comment\"># Media filtering</span>\n    exclude_external_images=<span class=\"hljs-literal\">True</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Explanation</strong>:</p>\n<ul>\n<li><strong><code>word_count_threshold</code></strong>: Ignores text blocks under X words. Helps skip trivial blocks like short nav or disclaimers.  </li>\n<li><strong><code>excluded_tags</code></strong>: Removes entire tags (<code>&lt;form&gt;</code>, <code>&lt;header&gt;</code>, <code>&lt;footer&gt;</code>, etc.).  </li>\n<li><strong>Link Filtering</strong>:  </li>\n<li><code>exclude_external_links</code>: Strips out external links and may remove them from <code>result.links</code>.  </li>\n<li><code>exclude_social_media_links</code>: Removes links pointing to known social media domains.  </li>\n<li><code>exclude_domains</code>: A custom list of domains to block if discovered in links.  </li>\n<li><code>exclude_social_media_domains</code>: A curated list (override or add to it) for social media sites.  </li>\n<li><strong>Media Filtering</strong>:  </li>\n<li><code>exclude_external_images</code>: Discards images not hosted on the same domain as the main page (or its subdomains).</li>\n</ul>\n<p>By default in case you set <code>exclude_social_media_links=True</code>, the following social media domains are excluded:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">[\n    <span class=\"hljs-string\">'facebook.com'</span>,\n    <span class=\"hljs-string\">'twitter.com'</span>,\n    <span class=\"hljs-string\">'x.com'</span>,\n    <span class=\"hljs-string\">'linkedin.com'</span>,\n    <span class=\"hljs-string\">'instagram.com'</span>,\n    <span class=\"hljs-string\">'pinterest.com'</span>,\n    <span class=\"hljs-string\">'tiktok.com'</span>,\n    <span class=\"hljs-string\">'snapchat.com'</span>,\n    <span class=\"hljs-string\">'reddit.com'</span>,\n]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"22-example-usage\">2.2 Example Usage</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        css_selector=<span class=\"hljs-string\">\"main.content\"</span>, \n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"ads.com\"</span>, <span class=\"hljs-string\">\"spammytrackers.net\"</span>],\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Cleaned HTML length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Note</strong>: If these parameters remove too much, reduce or disable them accordingly.</p>\n<hr>\n<h2 id=\"3-handling-iframes\">3. Handling Iframes</h2>\n<p>Some sites embed content in <code>&lt;iframe&gt;</code> tags. If you want that inline:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    <span class=\"hljs-comment\"># Merge iframe content into the final output</span>\n    process_iframes<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    remove_overlay_elements<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-comment\"># Remove GDPR/cookie consent popups (OneTrust, Cookiebot, etc.)</span>\n    remove_consent_popups<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Usage</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        process_iframes=<span class=\"hljs-literal\">True</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.org/iframe-demo\"</span>, \n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Iframe-merged length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"31-flattening-shadow-dom\">3.1 Flattening Shadow DOM</h2>\n<p>Sites built with <strong>Web Components</strong> (Stencil, Lit, Shoelace, Angular Elements, etc.) render content inside <a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Web_components/Using_shadow_DOM\">Shadow DOM</a> — an encapsulated sub-tree that is invisible to normal page serialization. The browser renders it on screen, but <code>page.content()</code> never includes it.</p>\n<p>Set <code>flatten_shadow_dom=True</code> to walk all shadow trees, resolve <code>&lt;slot&gt;</code> projections, and produce a single flat HTML document:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    <span class=\"hljs-comment\"># Flatten shadow DOM into the main document</span>\n    flatten_shadow_dom<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-comment\"># Give web components time to hydrate</span>\n    wait_until<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"load\"</span>,\n    delay_before_return_html<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3.0</span>,\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Full example</strong> — crawling a product page where specs live inside shadow roots:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        flatten_shadow_dom=<span class=\"hljs-literal\">True</span>,\n        wait_until=<span class=\"hljs-string\">\"load\"</span>,\n        delay_before_return_html=<span class=\"hljs-number\">3.0</span>,\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://store.boschrexroth.com/en/us/p/hydraulic-cylinder-r900999011\"</span>,\n            config=config,\n        )\n        <span class=\"hljs-comment\"># Without flatten_shadow_dom: ~1 KB of markdown (breadcrumbs only)</span>\n        <span class=\"hljs-comment\"># With flatten_shadow_dom:   ~33 KB (full product specs, downloads, etc.)</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>When <code>flatten_shadow_dom=True</code> is set, Crawl4AI also injects an init script that force-opens <strong>closed</strong> shadow roots (by patching <code>Element.prototype.attachShadow</code>), so even components that use <code>mode: 'closed'</code> become accessible.</p>\n<blockquote>\n<p><strong>Tip</strong>: Web components need JavaScript to run before they render content (a process called <em>hydration</em>). Use <code>wait_until=\"load\"</code> and a <code>delay_before_return_html</code> of 2–5 seconds to ensure components are fully hydrated before flattening.</p>\n</blockquote>\n<p>For a complete runnable example, see <a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/shadow_dom_crawling.py\"><code>shadow_dom_crawling.py</code></a>.</p>\n<hr>\n<h2 id=\"4-structured-extraction-examples\">4. Structured Extraction Examples</h2>\n<p>You can combine content selection with a more advanced extraction strategy. For instance, a <strong>CSS-based</strong> or <strong>LLM-based</strong> extraction strategy can run on the filtered HTML.</p>\n<h3 id=\"41-pattern-based-with-jsoncssextractionstrategy\">4.1 Pattern-Based with <code>JsonCssExtractionStrategy</code></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Minimal schema for repeated items</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"News Items\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"tr.athing\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.titleline a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, \n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.titleline a\"</span>, \n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, \n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>\n            }\n        ]\n    }\n\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Content filtering</span>\n        excluded_tags=[<span class=\"hljs-string\">\"form\"</span>, <span class=\"hljs-string\">\"header\"</span>],\n        exclude_domains=[<span class=\"hljs-string\">\"adsite.com\"</span>],\n\n        <span class=\"hljs-comment\"># CSS selection or entire page</span>\n        css_selector=<span class=\"hljs-string\">\"table.itemlist\"</span>,\n\n        <span class=\"hljs-comment\"># No caching for demonstration</span>\n        cache_mode=CacheMode.BYPASS,\n\n        <span class=\"hljs-comment\"># Extraction strategy</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema)\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>, \n            config=config\n        )\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Sample extracted item:\"</span>, data[:<span class=\"hljs-number\">1</span>])  <span class=\"hljs-comment\"># Show first item</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"42-llm-based-extraction\">4.2 LLM-Based Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">ArticleData</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    headline: <span class=\"hljs-built_in\">str</span>\n    summary: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    llm_strategy = LLMExtractionStrategy(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4\"</span>,api_token=<span class=\"hljs-string\">\"sk-YOUR_API_KEY\"</span>)\n        schema=ArticleData.schema(),\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract 'headline' and a short 'summary' from the content.\"</span>\n    )\n\n    config = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        word_count_threshold=<span class=\"hljs-number\">20</span>,\n        extraction_strategy=llm_strategy\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n        article = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(article)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Here, the crawler:</p>\n<ul>\n<li>Filters out external links (<code>exclude_external_links=True</code>).  </li>\n<li>Ignores very short text blocks (<code>word_count_threshold=20</code>).  </li>\n<li>Passes the final HTML to your LLM strategy for an AI-driven parse.</li>\n</ul>\n<hr>\n<h2 id=\"5-comprehensive-example\">5. Comprehensive Example</h2>\n<p>Below is a short function that unifies <strong>CSS selection</strong>, <strong>exclusion</strong> logic, and a pattern-based extraction, demonstrating how you can fine-tune your final data:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_main_articles</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span></span>):\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"ArticleBlock\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.article-block\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"headline\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"summary\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".summary\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"metadata\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"nested\"</span>,\n                <span class=\"hljs-string\">\"fields\"</span>: [\n                    {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"author\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".author\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n                    {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"date\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".date\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>}\n                ]\n            }\n        ]\n    }\n\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Keep only #main-content</span>\n        css_selector=<span class=\"hljs-string\">\"#main-content\"</span>,\n\n        <span class=\"hljs-comment\"># Filtering</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>],  \n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"somebadsite.com\"</span>],\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Extraction</span>\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=config)\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">None</span>\n        <span class=\"hljs-keyword\">return</span> json.loads(result.extracted_content)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    articles = <span class=\"hljs-keyword\">await</span> extract_main_articles(<span class=\"hljs-string\">\"https://news.ycombinator.com/newest\"</span>)\n    <span class=\"hljs-keyword\">if</span> articles:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted Articles:\"</span>, articles[:<span class=\"hljs-number\">2</span>])  <span class=\"hljs-comment\"># Show first 2</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Why This Works</strong>:\n- <strong>CSS</strong> scoping with <code>#main-content</code>.<br>\n- Multiple <strong>exclude_</strong> parameters to remove domains, external images, etc.<br>\n- A <strong>JsonCssExtractionStrategy</strong> to parse repeated article blocks.</p>\n<hr>\n<h2 id=\"6-scraping-modes\">6. Scraping Modes</h2>\n<p>Crawl4AI uses <code>LXMLWebScrapingStrategy</code> (LXML-based) as the default scraping strategy for HTML content processing. This strategy offers excellent performance, especially for large HTML documents.</p>\n<p><strong>Note:</strong> For backward compatibility, <code>WebScrapingStrategy</code> is still available as an alias for <code>LXMLWebScrapingStrategy</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LXMLWebScrapingStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Default configuration already uses LXMLWebScrapingStrategy</span>\n    config = CrawlerRunConfig()\n\n    <span class=\"hljs-comment\"># Or explicitly specify it if desired</span>\n    config_explicit = CrawlerRunConfig(\n        scraping_strategy=LXMLWebScrapingStrategy()\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>, \n            config=config\n        )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>You can also create your own custom scraping strategy by inheriting from <code>ContentScrapingStrategy</code>. The strategy must return a <code>ScrapingResult</code> object with the following structure:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> ContentScrapingStrategy, ScrapingResult, MediaItem, Media, Link, Links\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CustomScrapingStrategy</span>(<span class=\"hljs-title class_ inherited__\">ContentScrapingStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">scrap</span>(<span class=\"hljs-params\">self, url: <span class=\"hljs-built_in\">str</span>, html: <span class=\"hljs-built_in\">str</span>, **kwargs</span>) -&gt; ScrapingResult:\n        <span class=\"hljs-comment\"># Implement your custom scraping logic here</span>\n        <span class=\"hljs-keyword\">return</span> ScrapingResult(\n            cleaned_html=<span class=\"hljs-string\">\"&lt;html&gt;...&lt;/html&gt;\"</span>,  <span class=\"hljs-comment\"># Cleaned HTML content</span>\n            success=<span class=\"hljs-literal\">True</span>,                     <span class=\"hljs-comment\"># Whether scraping was successful</span>\n            media=Media(\n                images=[                      <span class=\"hljs-comment\"># List of images found</span>\n                    MediaItem(\n                        src=<span class=\"hljs-string\">\"https://example.com/image.jpg\"</span>,\n                        alt=<span class=\"hljs-string\">\"Image description\"</span>,\n                        desc=<span class=\"hljs-string\">\"Surrounding text\"</span>,\n                        score=<span class=\"hljs-number\">1</span>,\n                        <span class=\"hljs-built_in\">type</span>=<span class=\"hljs-string\">\"image\"</span>,\n                        group_id=<span class=\"hljs-number\">1</span>,\n                        <span class=\"hljs-built_in\">format</span>=<span class=\"hljs-string\">\"jpg\"</span>,\n                        width=<span class=\"hljs-number\">800</span>\n                    )\n                ],\n                videos=[],                    <span class=\"hljs-comment\"># List of videos (same structure as images)</span>\n                audios=[]                     <span class=\"hljs-comment\"># List of audio files (same structure as images)</span>\n            ),\n            links=Links(\n                internal=[                    <span class=\"hljs-comment\"># List of internal links</span>\n                    Link(\n                        href=<span class=\"hljs-string\">\"https://example.com/page\"</span>,\n                        text=<span class=\"hljs-string\">\"Link text\"</span>,\n                        title=<span class=\"hljs-string\">\"Link title\"</span>,\n                        base_domain=<span class=\"hljs-string\">\"example.com\"</span>\n                    )\n                ],\n                external=[]                   <span class=\"hljs-comment\"># List of external links (same structure)</span>\n            ),\n            metadata={                        <span class=\"hljs-comment\"># Additional metadata</span>\n                <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"Page Title\"</span>,\n                <span class=\"hljs-string\">\"description\"</span>: <span class=\"hljs-string\">\"Page description\"</span>\n            }\n        )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">ascrap</span>(<span class=\"hljs-params\">self, url: <span class=\"hljs-built_in\">str</span>, html: <span class=\"hljs-built_in\">str</span>, **kwargs</span>) -&gt; ScrapingResult:\n        <span class=\"hljs-comment\"># For simple cases, you can use the sync version</span>\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-keyword\">await</span> asyncio.to_thread(self.scrap, url, html, **kwargs)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"performance-considerations\">Performance Considerations</h3>\n<p>The LXML strategy provides excellent performance, particularly when processing large HTML documents, offering up to 10-20x faster processing compared to BeautifulSoup-based approaches.</p>\n<p>Benefits of LXML strategy:\n- Fast processing of large HTML documents (especially &gt;100KB)\n- Efficient memory usage\n- Good handling of well-formed HTML\n- Robust table detection and extraction</p>\n<h3 id=\"backward-compatibility\">Backward Compatibility</h3>\n<p>For users upgrading from earlier versions:\n- <code>WebScrapingStrategy</code> is now an alias for <code>LXMLWebScrapingStrategy</code>\n- Existing code using <code>WebScrapingStrategy</code> will continue to work without modification\n- No changes are required to your existing code</p>\n<hr>\n<h2 id=\"7-combining-css-selection-methods\">7. Combining CSS Selection Methods</h2>\n<p>You can combine <code>css_selector</code> and <code>target_elements</code> in powerful ways to achieve fine-grained control over your output:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Target specific content but preserve page context</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Focus markdown on main content and sidebar</span>\n        target_elements=[<span class=\"hljs-string\">\"#main-content\"</span>, <span class=\"hljs-string\">\".sidebar\"</span>],\n\n        <span class=\"hljs-comment\"># Global filters applied to entire page</span>\n        excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>, <span class=\"hljs-string\">\"header\"</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Use basic content thresholds</span>\n        word_count_threshold=<span class=\"hljs-number\">15</span>,\n\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/article\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Content focuses on specific elements, but all links still analyzed\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal links: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">'internal'</span>, []))}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"External links: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">'external'</span>, []))}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This approach gives you the best of both worlds:\n- Markdown generation and content extraction focus on the elements you care about\n- Links, images and other page data still give you the full context of the page\n- Content filtering still applies globally</p>\n<h2 id=\"8-conclusion\">8. Conclusion</h2>\n<p>By mixing <strong>target_elements</strong> or <strong>css_selector</strong> scoping, <strong>content filtering</strong> parameters, and advanced <strong>extraction strategies</strong>, you can precisely <strong>choose</strong> which data to keep. Key parameters in <strong><code>CrawlerRunConfig</code></strong> for content selection include:</p>\n<ol>\n<li><strong><code>target_elements</code></strong> – Array of CSS selectors to focus markdown generation and data extraction, while preserving full page context for links and media.</li>\n<li><strong><code>css_selector</code></strong> – Basic scoping to an element or region for all extraction processes.  </li>\n<li><strong><code>word_count_threshold</code></strong> – Skip short blocks.  </li>\n<li><strong><code>excluded_tags</code></strong> – Remove entire HTML tags.  </li>\n<li><strong><code>exclude_external_links</code></strong>, <strong><code>exclude_social_media_links</code></strong>, <strong><code>exclude_domains</code></strong> – Filter out unwanted links or domains.  </li>\n<li><strong><code>exclude_external_images</code></strong> – Remove images from external sources.  </li>\n<li><strong><code>process_iframes</code></strong> – Merge iframe content if needed.  </li>\n</ol>\n<p>Combine these with structured extraction (CSS, LLM-based, or others) to build powerful crawls that yield exactly the content you want, from raw or cleaned HTML up to sophisticated JSON structures. For more detail, see <a href=\"../../api/parameters/\">Configuration Reference</a>. Enjoy curating your data to the max!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/crawler-result/",
    "title": "Crawler Result - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"crawl-result-and-output\">Crawl Result and Output</h1>\n<p>When you call <code>arun()</code> on a page, Crawl4AI returns a <strong><code>CrawlResult</code></strong> object containing everything you might need—raw HTML, a cleaned version, optional screenshots or PDFs, structured extraction results, and more. This document explains those fields and how they map to different output types.  </p>\n<hr>\n<h2 id=\"1-the-crawlresult-model\">1. The <code>CrawlResult</code> Model</h2>\n<p>Below is the core schema. Each field captures a different aspect of the crawl’s result:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">MarkdownGenerationResult</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    raw_markdown: <span class=\"hljs-built_in\">str</span>\n    markdown_with_citations: <span class=\"hljs-built_in\">str</span>\n    references_markdown: <span class=\"hljs-built_in\">str</span>\n    fit_markdown: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    fit_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CrawlResult</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    url: <span class=\"hljs-built_in\">str</span>\n    html: <span class=\"hljs-built_in\">str</span>\n    fit_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    success: <span class=\"hljs-built_in\">bool</span>\n    cleaned_html: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    media: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    links: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>]] = {}\n    downloaded_files: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-built_in\">str</span>]] = <span class=\"hljs-literal\">None</span>\n    js_execution_result: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">Any</span>]] = <span class=\"hljs-literal\">None</span>\n    screenshot: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    pdf: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">bytes</span>] = <span class=\"hljs-literal\">None</span>\n    mhtml: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    markdown: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">Union</span>[<span class=\"hljs-built_in\">str</span>, MarkdownGenerationResult]] = <span class=\"hljs-literal\">None</span>\n    extracted_content: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    metadata: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    error_message: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    session_id: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    response_headers: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">dict</span>] = <span class=\"hljs-literal\">None</span>\n    status_code: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">int</span>] = <span class=\"hljs-literal\">None</span>\n    ssl_certificate: <span class=\"hljs-type\">Optional</span>[SSLCertificate] = <span class=\"hljs-literal\">None</span>\n    dispatch_result: <span class=\"hljs-type\">Optional</span>[DispatchResult] = <span class=\"hljs-literal\">None</span>\n    redirected_url: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n    redirected_status_code: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-built_in\">int</span>] = <span class=\"hljs-literal\">None</span>\n    network_requests: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">Any</span>]]] = <span class=\"hljs-literal\">None</span>\n    console_messages: <span class=\"hljs-type\">Optional</span>[<span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-type\">Any</span>]]] = <span class=\"hljs-literal\">None</span>\n    tables: <span class=\"hljs-type\">List</span>[<span class=\"hljs-type\">Dict</span>] = Field(default_factory=<span class=\"hljs-built_in\">list</span>)\n\n    <span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Config</span>:\n        arbitrary_types_allowed = <span class=\"hljs-literal\">True</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"table-key-fields-in-crawlresult\">Table: Key Fields in <code>CrawlResult</code></h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Field (Name &amp; Type)</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>url (<code>str</code>)</strong></td>\n<td>The final or actual URL crawled (in case of redirects).</td>\n</tr>\n<tr>\n<td><strong>html (<code>str</code>)</strong></td>\n<td>Original, unmodified page HTML. Good for debugging or custom processing.</td>\n</tr>\n<tr>\n<td><strong>fit_html (<code>Optional[str]</code>)</strong></td>\n<td>Preprocessed HTML optimized for extraction and content filtering.</td>\n</tr>\n<tr>\n<td><strong>success (<code>bool</code>)</strong></td>\n<td><code>True</code> if the crawl completed without major errors, else <code>False</code>.</td>\n</tr>\n<tr>\n<td><strong>cleaned_html (<code>Optional[str]</code>)</strong></td>\n<td>Sanitized HTML with scripts/styles removed; can exclude tags if configured via <code>excluded_tags</code> etc.</td>\n</tr>\n<tr>\n<td><strong>media (<code>Dict[str, List[Dict]]</code>)</strong></td>\n<td>Extracted media info (images, audio, etc.), each with attributes like <code>src</code>, <code>alt</code>, <code>score</code>, etc.</td>\n</tr>\n<tr>\n<td><strong>links (<code>Dict[str, List[Dict]]</code>)</strong></td>\n<td>Extracted link data, split by <code>internal</code> and <code>external</code>. Each link usually has <code>href</code>, <code>text</code>, etc.</td>\n</tr>\n<tr>\n<td><strong>downloaded_files (<code>Optional[List[str]]</code>)</strong></td>\n<td>If <code>accept_downloads=True</code> in <code>BrowserConfig</code>, this lists the filepaths of saved downloads.</td>\n</tr>\n<tr>\n<td><strong>js_execution_result (<code>Optional[Dict[str, Any]]</code>)</strong></td>\n<td>Results from JavaScript execution during crawling.</td>\n</tr>\n<tr>\n<td><strong>screenshot (<code>Optional[str]</code>)</strong></td>\n<td>Screenshot of the page (base64-encoded) if <code>screenshot=True</code>.</td>\n</tr>\n<tr>\n<td><strong>pdf (<code>Optional[bytes]</code>)</strong></td>\n<td>PDF of the page if <code>pdf=True</code>.</td>\n</tr>\n<tr>\n<td><strong>mhtml (<code>Optional[str]</code>)</strong></td>\n<td>MHTML snapshot of the page if <code>capture_mhtml=True</code>. Contains the full page with all resources.</td>\n</tr>\n<tr>\n<td><strong>markdown (<code>Optional[str or MarkdownGenerationResult]</code>)</strong></td>\n<td>It holds a <code>MarkdownGenerationResult</code>. Over time, this will be consolidated into <code>markdown</code>. The generator can provide raw markdown, citations, references, and optionally <code>fit_markdown</code>.</td>\n</tr>\n<tr>\n<td><strong>extracted_content (<code>Optional[str]</code>)</strong></td>\n<td>The output of a structured extraction (CSS/LLM-based) stored as JSON string or other text.</td>\n</tr>\n<tr>\n<td><strong>metadata (<code>Optional[dict]</code>)</strong></td>\n<td>Additional info about the crawl or extracted data.</td>\n</tr>\n<tr>\n<td><strong>error_message (<code>Optional[str]</code>)</strong></td>\n<td>If <code>success=False</code>, contains a short description of what went wrong.</td>\n</tr>\n<tr>\n<td><strong>session_id (<code>Optional[str]</code>)</strong></td>\n<td>The ID of the session used for multi-page or persistent crawling.</td>\n</tr>\n<tr>\n<td><strong>response_headers (<code>Optional[dict]</code>)</strong></td>\n<td>HTTP response headers, if captured.</td>\n</tr>\n<tr>\n<td><strong>status_code (<code>Optional[int]</code>)</strong></td>\n<td>HTTP status code (e.g., 200 for OK).</td>\n</tr>\n<tr>\n<td><strong>ssl_certificate (<code>Optional[SSLCertificate]</code>)</strong></td>\n<td>SSL certificate info if <code>fetch_ssl_certificate=True</code>.</td>\n</tr>\n<tr>\n<td><strong>dispatch_result (<code>Optional[DispatchResult]</code>)</strong></td>\n<td>Additional concurrency and resource usage information when crawling URLs in parallel.</td>\n</tr>\n<tr>\n<td><strong>redirected_url (<code>Optional[str]</code>)</strong></td>\n<td>The URL after any redirects (different from <code>url</code> which is the final URL).</td>\n</tr>\n<tr>\n<td><strong>redirected_status_code (<code>Optional[int]</code>)</strong></td>\n<td>HTTP status code of the final redirect destination (e.g., 200). <code>None</code> for non-HTTP requests (raw HTML, local files).</td>\n</tr>\n<tr>\n<td><strong>network_requests (<code>Optional[List[Dict[str, Any]]]</code>)</strong></td>\n<td>List of network requests, responses, and failures captured during the crawl if <code>capture_network_requests=True</code>.</td>\n</tr>\n<tr>\n<td><strong>console_messages (<code>Optional[List[Dict[str, Any]]]</code>)</strong></td>\n<td>List of browser console messages captured during the crawl if <code>capture_console_messages=True</code>.</td>\n</tr>\n<tr>\n<td><strong>tables (<code>List[Dict]</code>)</strong></td>\n<td>Table data extracted from HTML tables with structure <code>[{headers, rows, caption, summary}]</code>.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2 id=\"2-html-variants\">2. HTML Variants</h2>\n<h3 id=\"html-raw-html\"><code>html</code>: Raw HTML</h3>\n<p>Crawl4AI preserves the exact HTML as <code>result.html</code>. Useful for:</p>\n<ul>\n<li>Debugging page issues or checking the original content.</li>\n<li>Performing your own specialized parse if needed.</li>\n</ul>\n<h3 id=\"cleaned_html-sanitized\"><code>cleaned_html</code>: Sanitized</h3>\n<p>If you specify any cleanup or exclusion parameters in <code>CrawlerRunConfig</code> (like <code>excluded_tags</code>, <code>remove_forms</code>, etc.), you’ll see the result here:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-built_in\">config</span> = CrawlerRunConfig(\n    excluded_tags=[<span class=\"hljs-string\">\"form\"</span>, <span class=\"hljs-string\">\"header\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n    keep_data_attributes=False\n)\nresult = await crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, <span class=\"hljs-built_in\">config</span>=<span class=\"hljs-built_in\">config</span>)\n<span class=\"hljs-built_in\">print</span>(result.cleaned_html)  # Freed of forms, header, footer, data-* attributes\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"3-markdown-generation\">3. Markdown Generation</h2>\n<h3 id=\"31-markdown\">3.1 <code>markdown</code></h3>\n<ul>\n<li><strong><code>markdown</code></strong>: The current location for detailed markdown output, returning a <strong><code>MarkdownGenerationResult</code></strong> object.  </li>\n<li><strong><code>markdown_v2</code></strong>: Removed in v0.5. Accessing it now raises <code>AttributeError</code>; use <code>markdown</code>.</li>\n</ul>\n<p><strong><code>MarkdownGenerationResult</code></strong> Fields:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Field</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>raw_markdown</strong></td>\n<td>The basic HTML→Markdown conversion.</td>\n</tr>\n<tr>\n<td><strong>markdown_with_citations</strong></td>\n<td>Markdown including inline citations that reference links at the end.</td>\n</tr>\n<tr>\n<td><strong>references_markdown</strong></td>\n<td>The references/citations themselves (if <code>citations=True</code>).</td>\n</tr>\n<tr>\n<td><strong>fit_markdown</strong></td>\n<td>The filtered/“fit” markdown if a content filter was used.</td>\n</tr>\n<tr>\n<td><strong>fit_html</strong></td>\n<td>The filtered HTML that generated <code>fit_markdown</code>.</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"32-basic-example-with-a-markdown-generator\">3.2 Basic Example with a Markdown Generator</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\nconfig = CrawlerRunConfig(\n    markdown_generator=DefaultMarkdownGenerator(\n        options={<span class=\"hljs-string\">\"citations\"</span>: <span class=\"hljs-literal\">True</span>, <span class=\"hljs-string\">\"body_width\"</span>: <span class=\"hljs-number\">80</span>}  <span class=\"hljs-comment\"># e.g. pass html2text style options</span>\n    )\n)\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\nmd_res = result.markdown  <span class=\"hljs-comment\"># or eventually 'result.markdown'</span>\n<span class=\"hljs-built_in\">print</span>(md_res.raw_markdown[:<span class=\"hljs-number\">500</span>])\n<span class=\"hljs-built_in\">print</span>(md_res.markdown_with_citations)\n<span class=\"hljs-built_in\">print</span>(md_res.references_markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Note</strong>: If you use a filter like <code>PruningContentFilter</code>, you’ll get <code>fit_markdown</code> and <code>fit_html</code> as well.</p>\n<hr>\n<h2 id=\"4-structured-extraction-extracted_content\">4. Structured Extraction: <code>extracted_content</code></h2>\n<p>If you run a JSON-based extraction strategy (CSS, XPath, LLM, etc.), the structured data is <strong>not</strong> stored in <code>markdown</code>—it’s placed in <strong><code>result.extracted_content</code></strong> as a JSON string (or sometimes plain text).</p>\n<h3 id=\"example-css-extraction-with-raw-html\">Example: CSS Extraction with <code>raw://</code> HTML</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Example Items\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.item\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n    raw_html = <span class=\"hljs-string\">\"&lt;div class='item'&gt;&lt;h2&gt;Item 1&lt;/h2&gt;&lt;a href='https://example.com/item1'&gt;Link 1&lt;/a&gt;&lt;/div&gt;\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"raw://\"</span> + raw_html,\n            config=CrawlerRunConfig(\n                cache_mode=CacheMode.BYPASS,\n                extraction_strategy=JsonCssExtractionStrategy(schema)\n            )\n        )\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(data)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Here:\n- <code>url=\"raw://...\"</code> passes the HTML content directly, no network requests.<br>\n- The <strong>CSS</strong> extraction strategy populates <code>result.extracted_content</code> with the JSON array <code>[{\"title\": \"...\", \"link\": \"...\"}]</code>.</p>\n<hr>\n<h2 id=\"5-more-fields-links-media-tables-and-more\">5. More Fields: Links, Media, Tables and More</h2>\n<h3 id=\"51-links\">5.1 <code>links</code></h3>\n<p>A dictionary, typically with <code>\"internal\"</code> and <code>\"external\"</code> lists. Each entry might have <code>href</code>, <code>text</code>, <code>title</code>, etc. This is automatically captured if you haven’t disabled link extraction.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-built_in\">print</span>(result.links[<span class=\"hljs-string\">\"internal\"</span>][:3])  <span class=\"hljs-comment\"># Show first 3 internal links</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"52-media\">5.2 <code>media</code></h3>\n<p>Similarly, a dictionary with <code>\"images\"</code>, <code>\"audio\"</code>, <code>\"video\"</code>, etc. Each item could include <code>src</code>, <code>alt</code>, <code>score</code>, and more, if your crawler is set to gather them.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">images = result.media.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"images\"</span>, [])\n<span class=\"hljs-keyword\">for</span> img <span class=\"hljs-keyword\">in</span> images:\n    print(<span class=\"hljs-string\">\"Image URL:\"</span>, img[<span class=\"hljs-string\">\"src\"</span>], <span class=\"hljs-string\">\"Alt:\"</span>, img.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">\"alt\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"53-tables\">5.3 <code>tables</code></h3>\n<p>The <code>tables</code> field contains structured data extracted from HTML tables found on the crawled page. Tables are analyzed based on various criteria to determine if they are actual data tables (as opposed to layout tables), including:</p>\n<ul>\n<li>Presence of thead and tbody sections</li>\n<li>Use of th elements for headers</li>\n<li>Column consistency</li>\n<li>Text density</li>\n<li>And other factors</li>\n</ul>\n<p>Tables that score above the threshold (default: 7) are extracted and stored in result.tables.</p>\n<h3 id=\"accessing-table-data\">Accessing Table data:</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.w3schools.com/html/html_tables.asp\"</span>,\n            config=CrawlerRunConfig(\n                table_score_threshold=<span class=\"hljs-number\">7</span>  <span class=\"hljs-comment\"># Minimum score for table detection</span>\n            )\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.tables:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.tables)}</span> tables\"</span>)\n\n            <span class=\"hljs-keyword\">for</span> i, table <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(result.tables):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\nTable <span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>:\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Caption: <span class=\"hljs-subst\">{table.get(<span class=\"hljs-string\">'caption'</span>, <span class=\"hljs-string\">'No caption'</span>)}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Headers: <span class=\"hljs-subst\">{table[<span class=\"hljs-string\">'headers'</span>]}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Rows: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(table[<span class=\"hljs-string\">'rows'</span>])}</span>\"</span>)\n\n                <span class=\"hljs-comment\"># Print first few rows as example</span>\n                <span class=\"hljs-keyword\">for</span> j, row <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(table[<span class=\"hljs-string\">'rows'</span>][:<span class=\"hljs-number\">3</span>]):\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Row <span class=\"hljs-subst\">{j+<span class=\"hljs-number\">1</span>}</span>: <span class=\"hljs-subst\">{row}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"configuring-table-extraction\">Configuring Table Extraction:</h3>\n<p>You can adjust the sensitivity of the table detection algorithm with:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    table_score_threshold=5  <span class=\"hljs-comment\"># Lower value = more tables detected (default: 7)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Each extracted table contains: </p>\n<ul>\n<li><code>headers</code>: Column header names </li>\n<li><code>rows</code>: List of rows, each containing cell values</li>\n<li><code>caption</code>: Table caption text (if available) </li>\n<li><code>summary</code>: Table summary attribute (if specified)</li>\n</ul>\n<h3 id=\"table-extraction-tips\">Table Extraction Tips</h3>\n<ul>\n<li>Not all HTML tables are extracted - only those detected as \"data tables\" vs. layout tables.</li>\n<li>Tables with inconsistent cell counts, nested tables, or those used purely for layout may be skipped.</li>\n<li>If you're missing tables, try adjusting the <code>table_score_threshold</code> to a lower value (default is 7).</li>\n</ul>\n<p>The table detection algorithm scores tables based on features like consistent columns, presence of headers, text density, and more. Tables scoring above the threshold are considered data tables worth extracting.</p>\n<h3 id=\"54-screenshot-pdf-and-mhtml\">5.4 <code>screenshot</code>, <code>pdf</code>, and <code>mhtml</code></h3>\n<p>If you set <code>screenshot=True</code>, <code>pdf=True</code>, or <code>capture_mhtml=True</code> in <strong><code>CrawlerRunConfig</code></strong>, then:</p>\n<ul>\n<li><code>result.screenshot</code> contains a base64-encoded PNG string.</li>\n<li><code>result.pdf</code> contains raw PDF bytes (you can write them to a file).</li>\n<li><code>result.mhtml</code> contains the MHTML snapshot of the page as a string (you can write it to a .mhtml file).</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Save the PDF</span>\n<span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.pdf\"</span>, <span class=\"hljs-string\">\"wb\"</span>) <span class=\"hljs-keyword\">as</span> f:\n    f.write(result.pdf)\n\n<span class=\"hljs-comment\"># Save the MHTML</span>\n<span class=\"hljs-keyword\">if</span> result.mhtml:\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"page.mhtml\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        f.write(result.mhtml)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The MHTML (MIME HTML) format is particularly useful as it captures the entire web page including all of its resources (CSS, images, scripts, etc.) in a single file, making it perfect for archiving or offline viewing.</p>\n<h3 id=\"55-ssl_certificate\">5.5 <code>ssl_certificate</code></h3>\n<p>If <code>fetch_ssl_certificate=True</code>, <code>result.ssl_certificate</code> holds details about the site’s SSL cert, such as issuer, validity dates, etc.</p>\n<hr>\n<h2 id=\"6-accessing-these-fields\">6. Accessing These Fields</h2>\n<p>After you run:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">result</span> = await crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=some_config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Check any field:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">if result<span class=\"hljs-selector-class\">.success</span>:\n    <span class=\"hljs-built_in\">print</span>(result.status_code, result.response_headers)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Links found:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.<span class=\"hljs-built_in\">get</span>(<span class=\"hljs-string\">\"internal\"</span>, [])))\n    if result.markdown:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown snippet:\"</span>, result.markdown.raw_markdown[:<span class=\"hljs-number\">200</span>])\n    if result.extracted_content:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Structured JSON:\"</span>, result.extracted_content)\nelse:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Deprecation</strong>: Since v0.5 <code>markdown_v2</code>, <code>fit_markdown</code>, and <code>fit_html</code> are removed from <code>CrawlResult</code>. Use <code>result.markdown</code> for markdown output. It holds <code>MarkdownGenerationResult</code>, including <code>fit_html</code> and <code>fit_markdown</code>.</p>\n<hr>\n<h2 id=\"7-next-steps\">7. Next Steps</h2>\n<ul>\n<li><strong>Markdown Generation</strong>: Dive deeper into how to configure <code>DefaultMarkdownGenerator</code> and various filters.  </li>\n<li><strong>Content Filtering</strong>: Learn how to use <code>BM25ContentFilter</code> and <code>PruningContentFilter</code>.</li>\n<li><strong>Session &amp; Hooks</strong>: If you want to manipulate the page or preserve state across multiple <code>arun()</code> calls, see the hooking or session docs.  </li>\n<li><strong>LLM Extraction</strong>: For complex or unstructured content requiring AI-driven parsing, check the LLM-based strategies doc.</li>\n</ul>\n<p><strong>Enjoy</strong> exploring all that <code>CrawlResult</code> offers—whether you need raw HTML, sanitized output, markdown, or fully structured data, Crawl4AI has you covered!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/deep-crawling/",
    "title": "Deep Crawling - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"deep-crawling\">Deep Crawling</h1>\n<p>One of Crawl4AI's most powerful features is its ability to perform <strong>configurable deep crawling</strong> that can explore websites beyond a single page. With fine-tuned control over crawl depth, domain boundaries, and content filtering, Crawl4AI gives you the tools to extract precisely the content you need.</p>\n<p>In this tutorial, you'll learn:</p>\n<ol>\n<li>How to set up a <strong>Basic Deep Crawler</strong> with BFS strategy</li>\n<li>Understanding the difference between <strong>streamed and non-streamed</strong> output</li>\n<li>Implementing <strong>filters and scorers</strong> to target specific content</li>\n<li>Creating <strong>advanced filtering chains</strong> for sophisticated crawls</li>\n<li>Using <strong>BestFirstCrawling</strong> for intelligent exploration prioritization</li>\n<li><strong>Crash recovery</strong> for long-running production crawls</li>\n<li><strong>Prefetch mode</strong> for fast URL discovery  </li>\n</ol>\n<blockquote>\n<p><strong>Prerequisites</strong><br>\n- You’ve completed or read <a href=\"../simple-crawling/\">AsyncWebCrawler Basics</a> to understand how to run a simple crawl.<br>\n- You know how to configure <code>CrawlerRunConfig</code>.</p>\n</blockquote>\n<hr>\n<h2 id=\"1-quick-example\">1. Quick Example</h2>\n<p>Here's a minimal code snippet that implements a basic deep crawl using the <strong>BFSDeepCrawlStrategy</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_scraping_strategy <span class=\"hljs-keyword\">import</span> LXMLWebScrapingStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Configure a 2-level deep crawl</span>\n    config = CrawlerRunConfig(\n        deep_crawl_strategy=BFSDeepCrawlStrategy(\n            max_depth=<span class=\"hljs-number\">2</span>, \n            include_external=<span class=\"hljs-literal\">False</span>\n        ),\n        scraping_strategy=LXMLWebScrapingStrategy(),\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages in total\"</span>)\n\n        <span class=\"hljs-comment\"># Access individual results</span>\n        <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results[:<span class=\"hljs-number\">3</span>]:  <span class=\"hljs-comment\"># Show first 3 results</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"URL: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Depth: <span class=\"hljs-subst\">{result.metadata.get(<span class=\"hljs-string\">'depth'</span>, <span class=\"hljs-number\">0</span>)}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What's happening?</strong><br>\n- <code>BFSDeepCrawlStrategy(max_depth=2, include_external=False)</code> instructs Crawl4AI to:\n  - Crawl the starting page (depth 0) plus 2 more levels\n  - Stay within the same domain (don't follow external links)\n- Each result contains metadata like the crawl depth\n- Results are returned as a list after all crawling is complete</p>\n<hr>\n<h2 id=\"2-understanding-deep-crawling-strategy-options\">2. Understanding Deep Crawling Strategy Options</h2>\n<h3 id=\"21-bfsdeepcrawlstrategy-breadth-first-search\">2.1 BFSDeepCrawlStrategy (Breadth-First Search)</h3>\n<p>The <strong>BFSDeepCrawlStrategy</strong> uses a breadth-first approach, exploring all links at one depth before moving deeper:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\n<span class=\"hljs-comment\"># Basic configuration</span>\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">2</span>,               <span class=\"hljs-comment\"># Crawl initial page + 2 levels deep</span>\n    include_external=<span class=\"hljs-literal\">False</span>,    <span class=\"hljs-comment\"># Stay within the same domain</span>\n    max_pages=<span class=\"hljs-number\">50</span>,              <span class=\"hljs-comment\"># Maximum number of pages to crawl (optional)</span>\n    score_threshold=<span class=\"hljs-number\">0.3</span>,       <span class=\"hljs-comment\"># Minimum score for URLs to be crawled (optional)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key parameters:</strong>\n- <strong><code>max_depth</code></strong>: Number of levels to crawl beyond the starting page\n- <strong><code>include_external</code></strong>: Whether to follow links to other domains\n- <strong><code>max_pages</code></strong>: Maximum number of pages to crawl (default: infinite)\n- <strong><code>score_threshold</code></strong>: Minimum score for URLs to be crawled (default: -inf)\n- <strong><code>filter_chain</code></strong>: FilterChain instance for URL filtering\n- <strong><code>url_scorer</code></strong>: Scorer instance for evaluating URLs</p>\n<h3 id=\"22-dfsdeepcrawlstrategy-depth-first-search\">2.2 DFSDeepCrawlStrategy (Depth-First Search)</h3>\n<p>The <strong>DFSDeepCrawlStrategy</strong> uses a depth-first approach, explores as far down a branch as possible before backtracking.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> DFSDeepCrawlStrategy\n\n<span class=\"hljs-comment\"># Basic configuration</span>\nstrategy = DFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">2</span>,               <span class=\"hljs-comment\"># Crawl initial page + 2 levels deep</span>\n    include_external=<span class=\"hljs-literal\">False</span>,    <span class=\"hljs-comment\"># Stay within the same domain</span>\n    max_pages=<span class=\"hljs-number\">30</span>,              <span class=\"hljs-comment\"># Maximum number of pages to crawl (optional)</span>\n    score_threshold=<span class=\"hljs-number\">0.5</span>,       <span class=\"hljs-comment\"># Minimum score for URLs to be crawled (optional)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key parameters:</strong>\n- <strong><code>max_depth</code></strong>: Number of levels to crawl beyond the starting page\n- <strong><code>include_external</code></strong>: Whether to follow links to other domains\n- <strong><code>max_pages</code></strong>: Maximum number of pages to crawl (default: infinite)\n- <strong><code>score_threshold</code></strong>: Minimum score for URLs to be crawled (default: -inf)\n- <strong><code>filter_chain</code></strong>: FilterChain instance for URL filtering\n- <strong><code>url_scorer</code></strong>: Scorer instance for evaluating URLs</p>\n<h3 id=\"23-bestfirstcrawlingstrategy-recommended-deep-crawl-strategy\">2.3 BestFirstCrawlingStrategy (⭐️ - Recommended Deep crawl strategy)</h3>\n<p>For more intelligent crawling, use <strong>BestFirstCrawlingStrategy</strong> with scorers to prioritize the most relevant pages:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BestFirstCrawlingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling.scorers <span class=\"hljs-keyword\">import</span> KeywordRelevanceScorer\n\n<span class=\"hljs-comment\"># Create a scorer</span>\nscorer = KeywordRelevanceScorer(\n    keywords=[<span class=\"hljs-string\">\"crawl\"</span>, <span class=\"hljs-string\">\"example\"</span>, <span class=\"hljs-string\">\"async\"</span>, <span class=\"hljs-string\">\"configuration\"</span>],\n    weight=<span class=\"hljs-number\">0.7</span>\n)\n\n<span class=\"hljs-comment\"># Configure the strategy</span>\nstrategy = BestFirstCrawlingStrategy(\n    max_depth=<span class=\"hljs-number\">2</span>,\n    include_external=<span class=\"hljs-literal\">False</span>,\n    url_scorer=scorer,\n    max_pages=<span class=\"hljs-number\">25</span>,              <span class=\"hljs-comment\"># Maximum number of pages to crawl (optional)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This crawling approach:\n- Evaluates each discovered URL based on scorer criteria\n- Visits higher-scoring pages first\n- Helps focus crawl resources on the most relevant content\n- Can limit total pages crawled with <code>max_pages</code>\n- Does not need <code>score_threshold</code> as it naturally prioritizes by score</p>\n<hr>\n<h2 id=\"3-streaming-vs-non-streaming-results\">3. Streaming vs. Non-Streaming Results</h2>\n<p>Crawl4AI can return results in two modes:</p>\n<h3 id=\"31-non-streaming-mode-default\">3.1 Non-Streaming Mode (Default)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">config = CrawlerRunConfig(\n    deep_crawl_strategy=BFSDeepCrawlStrategy(max_depth=<span class=\"hljs-number\">1</span>),\n    stream=False  <span class=\"hljs-meta\"># Default behavior</span>\n)\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    # Wait <span class=\"hljs-keyword\">for</span> ALL results to be collected before returning\n    results</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n    <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n        process_result(result)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>When to use non-streaming mode:</strong>\n- You need the complete dataset before processing\n- You're performing batch operations on all results together\n- Crawl time isn't a critical factor</p>\n<h3 id=\"32-streaming-mode\">3.2 Streaming Mode</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    deep_crawl_strategy=BFSDeepCrawlStrategy(max_depth=<span class=\"hljs-number\">1</span>),\n    stream=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable streaming</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Returns an async iterator</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config):\n        <span class=\"hljs-comment\"># Process each result as it becomes available</span>\n        process_result(result)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Benefits of streaming mode:</strong>\n- Process results immediately as they're discovered\n- Start working with early results while crawling continues\n- Better for real-time applications or progressive display\n- Reduces memory pressure when handling many pages</p>\n<hr>\n<h2 id=\"4-filtering-content-with-filter-chains\">4. Filtering Content with Filter Chains</h2>\n<p>Filters help you narrow down which pages to crawl. Combine multiple filters using <strong>FilterChain</strong> for powerful targeting.</p>\n<h3 id=\"41-basic-url-pattern-filter\">4.1 Basic URL Pattern Filter</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling.filters import FilterChain, URLPatternFilter\n\n# <span class=\"hljs-keyword\">Only</span> follow URLs containing \"blog\" <span class=\"hljs-keyword\">or</span> \"docs\"\nurl_filter <span class=\"hljs-operator\">=</span> URLPatternFilter(patterns<span class=\"hljs-operator\">=</span>[\"*blog*\", \"*docs*\"])\n\nconfig <span class=\"hljs-operator\">=</span> CrawlerRunConfig(\n    deep_crawl_strategy<span class=\"hljs-operator\">=</span>BFSDeepCrawlStrategy(\n        max_depth<span class=\"hljs-operator\">=</span><span class=\"hljs-number\">1</span>,\n        filter_chain<span class=\"hljs-operator\">=</span>FilterChain([url_filter])\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"42-combining-multiple-filters\">4.2 Combining Multiple Filters</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai.deep_crawling.filters import (\n    FilterChain,\n    URLPatternFilter,\n    DomainFilter,\n    ContentTypeFilter\n)\n\n<span class=\"hljs-comment\"># Create a chain of filters</span>\nfilter_chain = FilterChain([\n    <span class=\"hljs-comment\"># Only follow URLs with specific patterns</span>\n    URLPatternFilter(patterns=[<span class=\"hljs-string\">\"*guide*\"</span>, <span class=\"hljs-string\">\"*tutorial*\"</span>]),\n\n    <span class=\"hljs-comment\"># Only crawl specific domains</span>\n    DomainFilter(\n        allowed_domains=[<span class=\"hljs-string\">\"docs.example.com\"</span>],\n        blocked_domains=[<span class=\"hljs-string\">\"old.docs.example.com\"</span>]\n    ),\n\n    <span class=\"hljs-comment\"># Only include specific content types</span>\n    ContentTypeFilter(allowed_types=[<span class=\"hljs-string\">\"text/html\"</span>])\n])\n\nconfig = CrawlerRunConfig(\n    deep_crawl_strategy=BFSDeepCrawlStrategy(\n        max_depth=2,\n        filter_chain=filter_chain\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"43-available-filter-types\">4.3 Available Filter Types</h3>\n<p>Crawl4AI includes several specialized filters:</p>\n<ul>\n<li><strong><code>URLPatternFilter</code></strong>: Matches URL patterns using wildcard syntax</li>\n<li><strong><code>DomainFilter</code></strong>: Controls which domains to include or exclude</li>\n<li><strong><code>ContentTypeFilter</code></strong>: Filters based on HTTP Content-Type</li>\n<li><strong><code>ContentRelevanceFilter</code></strong>: Uses similarity to a text query</li>\n<li><strong><code>SEOFilter</code></strong>: Evaluates SEO elements (meta tags, headers, etc.)</li>\n</ul>\n<hr>\n<h2 id=\"5-using-scorers-for-prioritized-crawling\">5. Using Scorers for Prioritized Crawling</h2>\n<p>Scorers assign priority values to discovered URLs, helping the crawler focus on the most relevant content first.</p>\n<h3 id=\"51-keywordrelevancescorer\">5.1 KeywordRelevanceScorer</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling.scorers <span class=\"hljs-keyword\">import</span> KeywordRelevanceScorer\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BestFirstCrawlingStrategy\n\n<span class=\"hljs-comment\"># Create a keyword relevance scorer</span>\nkeyword_scorer = KeywordRelevanceScorer(\n    keywords=[<span class=\"hljs-string\">\"crawl\"</span>, <span class=\"hljs-string\">\"example\"</span>, <span class=\"hljs-string\">\"async\"</span>, <span class=\"hljs-string\">\"configuration\"</span>],\n    weight=<span class=\"hljs-number\">0.7</span>  <span class=\"hljs-comment\"># Importance of this scorer (0.0 to 1.0)</span>\n)\n\nconfig = CrawlerRunConfig(\n    deep_crawl_strategy=BestFirstCrawlingStrategy(\n        max_depth=<span class=\"hljs-number\">2</span>,\n        url_scorer=keyword_scorer\n    ),\n    stream=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Recommended with BestFirstCrawling</span>\n)\n\n<span class=\"hljs-comment\"># Results will come in order of relevance score</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config):\n        score = result.metadata.get(<span class=\"hljs-string\">\"score\"</span>, <span class=\"hljs-number\">0</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Score: <span class=\"hljs-subst\">{score:<span class=\"hljs-number\">.2</span>f}</span> | <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>How scorers work:</strong>\n- Evaluate each discovered URL before crawling\n- Calculate relevance based on various signals\n- Help the crawler make intelligent choices about traversal order</p>\n<hr>\n<h2 id=\"6-advanced-filtering-techniques\">6. Advanced Filtering Techniques</h2>\n<h3 id=\"61-seo-filter-for-quality-assessment\">6.1 SEO Filter for Quality Assessment</h3>\n<p>The <strong>SEOFilter</strong> helps you identify pages with strong SEO characteristics:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai.deep_crawling.filters import FilterChain, SEOFilter\n\n<span class=\"hljs-comment\"># Create an SEO filter that looks for specific keywords in page metadata</span>\nseo_filter = SEOFilter(\n    threshold=0.5,  <span class=\"hljs-comment\"># Minimum score (0.0 to 1.0)</span>\n    keywords=[<span class=\"hljs-string\">\"tutorial\"</span>, <span class=\"hljs-string\">\"guide\"</span>, <span class=\"hljs-string\">\"documentation\"</span>]\n)\n\nconfig = CrawlerRunConfig(\n    deep_crawl_strategy=BFSDeepCrawlStrategy(\n        max_depth=1,\n        filter_chain=FilterChain([seo_filter])\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"62-content-relevance-filter\">6.2 Content Relevance Filter</h3>\n<p>The <strong>ContentRelevanceFilter</strong> analyzes the actual content of pages:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">from crawl4ai.deep_crawling.filters import FilterChain, ContentRelevanceFilter\n\n<span class=\"hljs-comment\"># Create a content relevance filter</span>\nrelevance_filter = ContentRelevanceFilter(\n    query=<span class=\"hljs-string\">\"Web crawling and data extraction with Python\"</span>,\n    threshold=0.7  <span class=\"hljs-comment\"># Minimum similarity score (0.0 to 1.0)</span>\n)\n\nconfig = CrawlerRunConfig(\n    deep_crawl_strategy=BFSDeepCrawlStrategy(\n        max_depth=1,\n        filter_chain=FilterChain([relevance_filter])\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This filter:\n- Measures semantic similarity between query and page content\n- It's a BM25-based relevance filter using head section content</p>\n<hr>\n<h2 id=\"7-building-a-complete-advanced-crawler\">7. Building a Complete Advanced Crawler</h2>\n<p>This example combines multiple techniques for a sophisticated crawl:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_scraping_strategy <span class=\"hljs-keyword\">import</span> LXMLWebScrapingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BestFirstCrawlingStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling.filters <span class=\"hljs-keyword\">import</span> (\n    FilterChain,\n    DomainFilter,\n    URLPatternFilter,\n    ContentTypeFilter\n)\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling.scorers <span class=\"hljs-keyword\">import</span> KeywordRelevanceScorer\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">run_advanced_crawler</span>():\n    <span class=\"hljs-comment\"># Create a sophisticated filter chain</span>\n    filter_chain = FilterChain([\n        <span class=\"hljs-comment\"># Domain boundaries</span>\n        DomainFilter(\n            allowed_domains=[<span class=\"hljs-string\">\"docs.example.com\"</span>],\n            blocked_domains=[<span class=\"hljs-string\">\"old.docs.example.com\"</span>]\n        ),\n\n        <span class=\"hljs-comment\"># URL patterns to include</span>\n        URLPatternFilter(patterns=[<span class=\"hljs-string\">\"*guide*\"</span>, <span class=\"hljs-string\">\"*tutorial*\"</span>, <span class=\"hljs-string\">\"*blog*\"</span>]),\n\n        <span class=\"hljs-comment\"># Content type filtering</span>\n        ContentTypeFilter(allowed_types=[<span class=\"hljs-string\">\"text/html\"</span>])\n    ])\n\n    <span class=\"hljs-comment\"># Create a relevance scorer</span>\n    keyword_scorer = KeywordRelevanceScorer(\n        keywords=[<span class=\"hljs-string\">\"crawl\"</span>, <span class=\"hljs-string\">\"example\"</span>, <span class=\"hljs-string\">\"async\"</span>, <span class=\"hljs-string\">\"configuration\"</span>],\n        weight=<span class=\"hljs-number\">0.7</span>\n    )\n\n    <span class=\"hljs-comment\"># Set up the configuration</span>\n    config = CrawlerRunConfig(\n        deep_crawl_strategy=BestFirstCrawlingStrategy(\n            max_depth=<span class=\"hljs-number\">2</span>,\n            include_external=<span class=\"hljs-literal\">False</span>,\n            filter_chain=filter_chain,\n            url_scorer=keyword_scorer\n        ),\n        scraping_strategy=LXMLWebScrapingStrategy(),\n        stream=<span class=\"hljs-literal\">True</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-comment\"># Execute the crawl</span>\n    results = []\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://docs.example.com\"</span>, config=config):\n            results.append(result)\n            score = result.metadata.get(<span class=\"hljs-string\">\"score\"</span>, <span class=\"hljs-number\">0</span>)\n            depth = result.metadata.get(<span class=\"hljs-string\">\"depth\"</span>, <span class=\"hljs-number\">0</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Depth: <span class=\"hljs-subst\">{depth}</span> | Score: <span class=\"hljs-subst\">{score:<span class=\"hljs-number\">.2</span>f}</span> | <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Analyze the results</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> high-value pages\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Average score: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">sum</span>(r.metadata.get(<span class=\"hljs-string\">'score'</span>, <span class=\"hljs-number\">0</span>) <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> results) / <span class=\"hljs-built_in\">len</span>(results):<span class=\"hljs-number\">.2</span>f}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Group by depth</span>\n    depth_counts = {}\n    <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n        depth = result.metadata.get(<span class=\"hljs-string\">\"depth\"</span>, <span class=\"hljs-number\">0</span>)\n        depth_counts[depth] = depth_counts.get(depth, <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Pages crawled by depth:\"</span>)\n    <span class=\"hljs-keyword\">for</span> depth, count <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">sorted</span>(depth_counts.items()):\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Depth <span class=\"hljs-subst\">{depth}</span>: <span class=\"hljs-subst\">{count}</span> pages\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(run_advanced_crawler())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"8-limiting-and-controlling-crawl-size\">8. Limiting and Controlling Crawl Size</h2>\n<h3 id=\"81-using-max_pages\">8.1 Using max_pages</h3>\n<p>You can limit the total number of pages crawled with the <code>max_pages</code> parameter:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Limit to exactly 20 pages regardless of depth</span>\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=3,\n    max_pages=20\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This feature is useful for:\n- Controlling API costs\n- Setting predictable execution times\n- Focusing on the most important content\n- Testing crawl configurations before full execution</p>\n<h3 id=\"82-using-score_threshold\">8.2 Using score_threshold</h3>\n<p>For BFS and DFS strategies, you can set a minimum score threshold to only crawl high-quality pages:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Only follow links with scores above 0.4</span>\nstrategy = DFSDeepCrawlStrategy(\n    max_depth=2,\n    url_scorer=KeywordRelevanceScorer(keywords=[<span class=\"hljs-string\">\"api\"</span>, <span class=\"hljs-string\">\"guide\"</span>, <span class=\"hljs-string\">\"reference\"</span>]),\n    score_threshold=0.4  <span class=\"hljs-comment\"># Skip URLs with scores below this value</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Note that for BestFirstCrawlingStrategy, score_threshold is not needed since pages are already processed in order of highest score first.</p>\n<h2 id=\"9-common-pitfalls-tips\">9. Common Pitfalls &amp; Tips</h2>\n<p>1.<strong>Set realistic limits.</strong> Be cautious with <code>max_depth</code> values &gt; 3, which can exponentially increase crawl size. Use <code>max_pages</code> to set hard limits.</p>\n<p>2.<strong>Don't neglect the scoring component.</strong> BestFirstCrawling works best with well-tuned scorers. Experiment with keyword weights for optimal prioritization.</p>\n<p>3.<strong>Be a good web citizen.</strong>  Respect robots.txt. (disabled by default)</p>\n<p>4.<strong>Handle page errors gracefully.</strong> Not all pages will be accessible. Check <code>result.status</code> when processing results.</p>\n<p>5.<strong>Balance breadth vs. depth.</strong> Choose your strategy wisely - BFS for comprehensive coverage, DFS for deep exploration, BestFirst for focused relevance-based crawling.</p>\n<p>6.<strong>Preserve HTTPS for security.</strong> If crawling HTTPS sites that redirect to HTTP, use <code>preserve_https_for_internal_links=True</code> to maintain secure connections:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    deep_crawl_strategy<span class=\"hljs-punctuation\">=</span>BFSDeepCrawlStrategy<span class=\"hljs-punctuation\">(</span>max_depth<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">2</span><span class=\"hljs-punctuation\">)</span>,\n    preserve_https_for_internal_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Keep HTTPS even if server redirects to HTTP</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This is especially useful for security-conscious crawling or when dealing with sites that support both protocols.</p>\n<hr>\n<h2 id=\"10-crash-recovery-for-long-running-crawls\">10. Crash Recovery for Long-Running Crawls</h2>\n<p>For production deployments, especially in cloud environments where instances can be terminated unexpectedly, Crawl4AI provides built-in crash recovery support for all deep crawl strategies.</p>\n<h3 id=\"101-enabling-state-persistence\">10.1 Enabling State Persistence</h3>\n<p>All deep crawl strategies (BFS, DFS, Best-First) support two optional parameters:</p>\n<ul>\n<li><strong><code>resume_state</code></strong>: Pass a previously saved state to resume from a checkpoint</li>\n<li><strong><code>on_state_change</code></strong>: Async callback fired after each URL is processed</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n<span class=\"hljs-keyword\">import</span> json\n\n<span class=\"hljs-comment\"># Callback to save state after each URL</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">save_state_to_redis</span>(<span class=\"hljs-params\">state: <span class=\"hljs-built_in\">dict</span></span>):\n    <span class=\"hljs-keyword\">await</span> redis.<span class=\"hljs-built_in\">set</span>(<span class=\"hljs-string\">\"crawl_state\"</span>, json.dumps(state))\n\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">3</span>,\n    on_state_change=save_state_to_redis,  <span class=\"hljs-comment\"># Called after each URL</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"102-state-structure\">10.2 State Structure</h3>\n<p>The state dictionary is JSON-serializable and contains:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">{\n    <span class=\"hljs-string\">\"strategy_type\"</span>: <span class=\"hljs-string\">\"bfs\"</span>,  <span class=\"hljs-comment\"># or \"dfs\", \"best_first\"</span>\n    <span class=\"hljs-string\">\"visited\"</span>: [<span class=\"hljs-string\">\"url1\"</span>, <span class=\"hljs-string\">\"url2\"</span>, ...],  <span class=\"hljs-comment\"># Already crawled URLs</span>\n    <span class=\"hljs-string\">\"pending\"</span>: [{<span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"...\"</span>, <span class=\"hljs-string\">\"parent_url\"</span>: <span class=\"hljs-string\">\"...\"</span>}],  <span class=\"hljs-comment\"># Queue/stack</span>\n    <span class=\"hljs-string\">\"depths\"</span>: {<span class=\"hljs-string\">\"url1\"</span>: 0, <span class=\"hljs-string\">\"url2\"</span>: 1},  <span class=\"hljs-comment\"># Depth tracking</span>\n    <span class=\"hljs-string\">\"pages_crawled\"</span>: 42  <span class=\"hljs-comment\"># Counter</span>\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"103-resuming-from-a-checkpoint\">10.3 Resuming from a Checkpoint</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\n<span class=\"hljs-comment\"># Load saved state (e.g., from Redis, database, or file)</span>\nsaved_state = json.loads(<span class=\"hljs-keyword\">await</span> redis.get(<span class=\"hljs-string\">\"crawl_state\"</span>))\n\n<span class=\"hljs-comment\"># Resume crawling from where we left off</span>\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">3</span>,\n    resume_state=saved_state,  <span class=\"hljs-comment\"># Continue from checkpoint</span>\n    on_state_change=save_state_to_redis,  <span class=\"hljs-comment\"># Keep saving progress</span>\n)\n\nconfig = CrawlerRunConfig(deep_crawl_strategy=strategy)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Will skip already-visited URLs and continue from pending queue</span>\n    results = <span class=\"hljs-keyword\">await</span> crawler.arun(start_url, config=config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"104-manual-state-export\">10.4 Manual State Export</h3>\n<p>You can export the last captured state using <code>export_state()</code>. Note that this requires <code>on_state_change</code> to be set (state is captured in the callback):</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n\ncaptured_state = <span class=\"hljs-literal\">None</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">capture_state</span>(<span class=\"hljs-params\">state: <span class=\"hljs-built_in\">dict</span></span>):\n    <span class=\"hljs-keyword\">global</span> captured_state\n    captured_state = state\n\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">2</span>,\n    on_state_change=capture_state,  <span class=\"hljs-comment\"># Required for state capture</span>\n)\nconfig = CrawlerRunConfig(deep_crawl_strategy=strategy)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    results = <span class=\"hljs-keyword\">await</span> crawler.arun(start_url, config=config)\n\n<span class=\"hljs-comment\"># Get the last captured state</span>\nstate = strategy.export_state()\n<span class=\"hljs-keyword\">if</span> state:\n    <span class=\"hljs-comment\"># Save to your preferred storage</span>\n    <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"crawl_checkpoint.json\"</span>, <span class=\"hljs-string\">\"w\"</span>) <span class=\"hljs-keyword\">as</span> f:\n        json.dump(state, f)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"105-complete-example-redis-based-recovery\">10.5 Complete Example: Redis-Based Recovery</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> redis.asyncio <span class=\"hljs-keyword\">as</span> redis\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\nREDIS_KEY = <span class=\"hljs-string\">\"crawl4ai:crawl_state\"</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    redis_client = redis.Redis(host=<span class=\"hljs-string\">'localhost'</span>, port=<span class=\"hljs-number\">6379</span>, db=<span class=\"hljs-number\">0</span>)\n\n    <span class=\"hljs-comment\"># Check for existing state</span>\n    saved_state = <span class=\"hljs-literal\">None</span>\n    existing = <span class=\"hljs-keyword\">await</span> redis_client.get(REDIS_KEY)\n    <span class=\"hljs-keyword\">if</span> existing:\n        saved_state = json.loads(existing)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Resuming from checkpoint: <span class=\"hljs-subst\">{saved_state[<span class=\"hljs-string\">'pages_crawled'</span>]}</span> pages already crawled\"</span>)\n\n    <span class=\"hljs-comment\"># State persistence callback</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">persist_state</span>(<span class=\"hljs-params\">state: <span class=\"hljs-built_in\">dict</span></span>):\n        <span class=\"hljs-keyword\">await</span> redis_client.<span class=\"hljs-built_in\">set</span>(REDIS_KEY, json.dumps(state))\n\n    <span class=\"hljs-comment\"># Create strategy with recovery support</span>\n    strategy = BFSDeepCrawlStrategy(\n        max_depth=<span class=\"hljs-number\">3</span>,\n        max_pages=<span class=\"hljs-number\">100</span>,\n        resume_state=saved_state,\n        on_state_change=persist_state,\n    )\n\n    config = CrawlerRunConfig(deep_crawl_strategy=strategy, stream=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-keyword\">try</span>:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n            <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n    <span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl interrupted: <span class=\"hljs-subst\">{e}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"State saved - restart to resume\"</span>)\n    <span class=\"hljs-keyword\">finally</span>:\n        <span class=\"hljs-keyword\">await</span> redis_client.close()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"106-zero-overhead\">10.6 Zero Overhead</h3>\n<p>When <code>resume_state=None</code> and <code>on_state_change=None</code> (the defaults), there is no performance impact. State tracking only activates when you enable these features.</p>\n<hr>\n<h2 id=\"11-cancellation-support-for-deep-crawls\">11. Cancellation Support for Deep Crawls</h2>\n<p>For production environments like cloud platforms, you often need to stop a running crawl mid-execution—whether the user changed their mind, specified the wrong URL, or wants to control costs. Crawl4AI provides built-in cancellation support for all deep crawl strategies.</p>\n<h3 id=\"111-two-ways-to-cancel\">11.1 Two Ways to Cancel</h3>\n<p><strong>Option A: Callback-based cancellation</strong> (recommended for external systems)</p>\n<p>Use <code>should_cancel</code> to check an external source (Redis, database, API) before each URL:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">check_if_cancelled</span>():\n    <span class=\"hljs-comment\"># Check Redis, database, or any external source</span>\n    job = <span class=\"hljs-keyword\">await</span> redis.get(<span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>\"</span>)\n    <span class=\"hljs-keyword\">return</span> job.get(<span class=\"hljs-string\">\"status\"</span>) == <span class=\"hljs-string\">\"cancelled\"</span>\n\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">3</span>,\n    max_pages=<span class=\"hljs-number\">1000</span>,\n    should_cancel=check_if_cancelled,  <span class=\"hljs-comment\"># Called before each URL</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Option B: Direct cancellation</strong> (for in-process control)</p>\n<p>Call <code>cancel()</code> directly on the strategy instance:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">strategy <span class=\"hljs-punctuation\">=</span> BFSDeepCrawlStrategy<span class=\"hljs-punctuation\">(</span>max_depth<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3</span>, max_pages<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">1000</span><span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># In another coroutine or thread:</span>\nstrategy.cancel<span class=\"hljs-punctuation\">(</span><span class=\"hljs-punctuation\">)</span>  <span class=\"hljs-comment\"># Thread-safe, stops before next URL</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"112-checking-cancellation-status\">11.2 Checking Cancellation Status</h3>\n<p>Use the <code>cancelled</code> property to check if a crawl was cancelled:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    results = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config=config)\n\n<span class=\"hljs-keyword\">if</span> strategy.cancelled:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl was cancelled after <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n<span class=\"hljs-keyword\">else</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl completed with <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"113-state-notifications-include-cancelled-flag\">11.3 State Notifications Include Cancelled Flag</h3>\n<p>When using <code>on_state_change</code>, the state dictionary includes a <code>cancelled</code> field:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">handle_state</span>(<span class=\"hljs-params\">state: <span class=\"hljs-built_in\">dict</span></span>):\n    <span class=\"hljs-keyword\">if</span> state.get(<span class=\"hljs-string\">\"cancelled\"</span>):\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl was cancelled!\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{state[<span class=\"hljs-string\">'pages_crawled'</span>]}</span> pages before cancellation\"</span>)\n    <span class=\"hljs-comment\"># Save state for potential resume</span>\n    <span class=\"hljs-keyword\">await</span> redis.<span class=\"hljs-built_in\">set</span>(<span class=\"hljs-string\">\"crawl_state\"</span>, json.dumps(state))\n\nstrategy = BFSDeepCrawlStrategy(\n    max_depth=<span class=\"hljs-number\">3</span>,\n    should_cancel=check_cancelled,\n    on_state_change=handle_state,\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"114-key-behaviors\">11.4 Key Behaviors</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Scenario</th>\n<th>Behavior</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Cancel before first URL</td>\n<td>Returns empty results, <code>cancelled=True</code></td>\n</tr>\n<tr>\n<td>Cancel during crawl</td>\n<td>Completes current URL, then stops</td>\n</tr>\n<tr>\n<td>Callback raises exception</td>\n<td>Logged as warning, crawl continues (fail-open)</td>\n</tr>\n<tr>\n<td>Strategy reuse after cancel</td>\n<td>Works normally (cancel flag auto-resets)</td>\n</tr>\n<tr>\n<td>Sync callback function</td>\n<td>Supported (auto-detected and handled)</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"115-complete-example-cloud-platform-job-cancellation\">11.5 Complete Example: Cloud Platform Job Cancellation</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> redis.asyncio <span class=\"hljs-keyword\">as</span> redis\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">run_cancellable_crawl</span>(<span class=\"hljs-params\">job_id: <span class=\"hljs-built_in\">str</span>, start_url: <span class=\"hljs-built_in\">str</span></span>):\n    redis_client = redis.Redis(host=<span class=\"hljs-string\">'localhost'</span>, port=<span class=\"hljs-number\">6379</span>, db=<span class=\"hljs-number\">0</span>)\n\n    <span class=\"hljs-comment\"># Check external cancellation source</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">check_cancelled</span>():\n        status = <span class=\"hljs-keyword\">await</span> redis_client.get(<span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>:status\"</span>)\n        <span class=\"hljs-keyword\">return</span> status == <span class=\"hljs-string\">b\"cancelled\"</span>\n\n    <span class=\"hljs-comment\"># Save progress for monitoring and recovery</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">save_progress</span>(<span class=\"hljs-params\">state: <span class=\"hljs-built_in\">dict</span></span>):\n        <span class=\"hljs-keyword\">await</span> redis_client.<span class=\"hljs-built_in\">set</span>(\n            <span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>:state\"</span>,\n            json.dumps(state)\n        )\n        <span class=\"hljs-comment\"># Update job progress</span>\n        <span class=\"hljs-keyword\">await</span> redis_client.<span class=\"hljs-built_in\">set</span>(\n            <span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>:pages_crawled\"</span>,\n            state[<span class=\"hljs-string\">\"pages_crawled\"</span>]\n        )\n\n    strategy = BFSDeepCrawlStrategy(\n        max_depth=<span class=\"hljs-number\">3</span>,\n        max_pages=<span class=\"hljs-number\">500</span>,\n        should_cancel=check_cancelled,\n        on_state_change=save_progress,\n    )\n\n    config = CrawlerRunConfig(\n        deep_crawl_strategy=strategy,\n        stream=<span class=\"hljs-literal\">True</span>,\n    )\n\n    results = []\n    <span class=\"hljs-keyword\">try</span>:\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n            <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun(start_url, config=config):\n                results.append(result)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n    <span class=\"hljs-keyword\">finally</span>:\n        <span class=\"hljs-comment\"># Report final status</span>\n        <span class=\"hljs-keyword\">if</span> strategy.cancelled:\n            <span class=\"hljs-keyword\">await</span> redis_client.<span class=\"hljs-built_in\">set</span>(<span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>:status\"</span>, <span class=\"hljs-string\">\"cancelled\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Job cancelled after <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-keyword\">await</span> redis_client.<span class=\"hljs-built_in\">set</span>(<span class=\"hljs-string\">f\"job:<span class=\"hljs-subst\">{job_id}</span>:status\"</span>, <span class=\"hljs-string\">\"completed\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Job completed with <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n\n        <span class=\"hljs-keyword\">await</span> redis_client.close()\n\n    <span class=\"hljs-keyword\">return</span> results\n\n<span class=\"hljs-comment\"># Usage</span>\n<span class=\"hljs-comment\"># asyncio.run(run_cancellable_crawl(\"job-123\", \"https://example.com\"))</span>\n<span class=\"hljs-comment\">#</span>\n<span class=\"hljs-comment\"># To cancel from another process:</span>\n<span class=\"hljs-comment\"># redis_client.set(\"job:job-123:status\", \"cancelled\")</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"116-supported-strategies\">11.6 Supported Strategies</h3>\n<p>Cancellation works identically across all deep crawl strategies:</p>\n<ul>\n<li><strong>BFSDeepCrawlStrategy</strong> - Breadth-first search</li>\n<li><strong>DFSDeepCrawlStrategy</strong> - Depth-first search</li>\n<li><strong>BestFirstCrawlingStrategy</strong> - Priority-based crawling</li>\n</ul>\n<p>All strategies support:\n- <code>should_cancel</code> callback parameter\n- <code>cancel()</code> method\n- <code>cancelled</code> property</p>\n<hr>\n<h2 id=\"12-prefetch-mode-for-fast-url-discovery\">12. Prefetch Mode for Fast URL Discovery</h2>\n<p>When you need to quickly discover URLs without full page processing, use <strong>prefetch mode</strong>. This is ideal for two-phase crawling where you first map the site, then selectively process specific pages.</p>\n<h3 id=\"121-enabling-prefetch-mode\">12.1 Enabling Prefetch Mode</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\nconfig = CrawlerRunConfig(prefetch=<span class=\"hljs-literal\">True</span>)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n    <span class=\"hljs-comment\"># Result contains only HTML and links - no markdown, no extraction</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links[<span class=\"hljs-string\">'internal'</span>])}</span> internal links\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.links[<span class=\"hljs-string\">'external'</span>])}</span> external links\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"122-what-gets-skipped\">12.2 What Gets Skipped</h3>\n<p>Prefetch mode uses a fast path that bypasses heavy processing:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Processing Step</th>\n<th>Normal Mode</th>\n<th>Prefetch Mode</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fetch HTML</td>\n<td>✅</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>Extract links</td>\n<td>✅</td>\n<td>✅ (fast <code>quick_extract_links()</code>)</td>\n</tr>\n<tr>\n<td>Generate markdown</td>\n<td>✅</td>\n<td>❌ Skipped</td>\n</tr>\n<tr>\n<td>Content scraping</td>\n<td>✅</td>\n<td>❌ Skipped</td>\n</tr>\n<tr>\n<td>Media extraction</td>\n<td>✅</td>\n<td>❌ Skipped</td>\n</tr>\n<tr>\n<td>LLM extraction</td>\n<td>✅</td>\n<td>❌ Skipped</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"123-performance-benefit\">12.3 Performance Benefit</h3>\n<ul>\n<li><strong>Normal mode</strong>: Full pipeline (~2-5 seconds per page)</li>\n<li><strong>Prefetch mode</strong>: HTML + links only (~200-500ms per page)</li>\n</ul>\n<p>This makes prefetch mode <strong>5-10x faster</strong> for URL discovery.</p>\n<h3 id=\"124-two-phase-crawling-pattern\">12.4 Two-Phase Crawling Pattern</h3>\n<p>The most common use case is two-phase crawling:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">two_phase_crawl</span>(<span class=\"hljs-params\">start_url: <span class=\"hljs-built_in\">str</span></span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># ═══════════════════════════════════════════════</span>\n        <span class=\"hljs-comment\"># Phase 1: Fast discovery (prefetch mode)</span>\n        <span class=\"hljs-comment\"># ═══════════════════════════════════════════════</span>\n        prefetch_config = CrawlerRunConfig(prefetch=<span class=\"hljs-literal\">True</span>)\n        discovery = <span class=\"hljs-keyword\">await</span> crawler.arun(start_url, config=prefetch_config)\n\n        all_urls = [link[<span class=\"hljs-string\">\"href\"</span>] <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> discovery.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])]\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Discovered <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(all_urls)}</span> URLs\"</span>)\n\n        <span class=\"hljs-comment\"># Filter to URLs you care about</span>\n        blog_urls = [url <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> all_urls <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"/blog/\"</span> <span class=\"hljs-keyword\">in</span> url]\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(blog_urls)}</span> blog posts to process\"</span>)\n\n        <span class=\"hljs-comment\"># ═══════════════════════════════════════════════</span>\n        <span class=\"hljs-comment\"># Phase 2: Full processing on selected URLs only</span>\n        <span class=\"hljs-comment\"># ═══════════════════════════════════════════════</span>\n        full_config = CrawlerRunConfig(\n            <span class=\"hljs-comment\"># Your normal extraction settings</span>\n            word_count_threshold=<span class=\"hljs-number\">100</span>,\n            remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n        )\n\n        results = []\n        <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> blog_urls:\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config=full_config)\n            <span class=\"hljs-keyword\">if</span> result.success:\n                results.append(result)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Processed: <span class=\"hljs-subst\">{url}</span>\"</span>)\n\n        <span class=\"hljs-keyword\">return</span> results\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    results = asyncio.run(two_phase_crawl(<span class=\"hljs-string\">\"https://example.com\"</span>))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Fully processed <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"125-use-cases\">12.5 Use Cases</h3>\n<ul>\n<li><strong>Site mapping</strong>: Quickly discover all URLs before deciding what to process</li>\n<li><strong>Link validation</strong>: Check which pages exist without heavy processing</li>\n<li><strong>Selective deep crawl</strong>: Prefetch to find URLs, filter by pattern, then full crawl</li>\n<li><strong>Crawl planning</strong>: Estimate crawl size before committing resources</li>\n</ul>\n<hr>\n<h2 id=\"13-summary-next-steps\">13. Summary &amp; Next Steps</h2>\n<p>In this <strong>Deep Crawling with Crawl4AI</strong> tutorial, you learned to:</p>\n<ul>\n<li>Configure <strong>BFSDeepCrawlStrategy</strong>, <strong>DFSDeepCrawlStrategy</strong>, and <strong>BestFirstCrawlingStrategy</strong></li>\n<li>Process results in streaming or non-streaming mode</li>\n<li>Apply filters to target specific content</li>\n<li>Use scorers to prioritize the most relevant pages</li>\n<li>Limit crawls with <code>max_pages</code> and <code>score_threshold</code> parameters</li>\n<li>Build a complete advanced crawler with combined techniques</li>\n<li><strong>Implement crash recovery</strong> with <code>resume_state</code> and <code>on_state_change</code> for production deployments</li>\n<li><strong>Cancel running crawls</strong> with <code>should_cancel</code> callback or <code>cancel()</code> method for cloud platform job management</li>\n<li><strong>Use prefetch mode</strong> for fast URL discovery and two-phase crawling</li>\n</ul>\n<p>With these tools, you can efficiently extract structured data from websites at scale, focusing precisely on the content you need for your specific use case.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/domain-mapping/",
    "title": "Domain Mapping - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"domain-mapping-discover-every-url-under-a-domain\">Domain Mapping: Discover Every URL Under a Domain</h1>\n<h2 id=\"what-is-domain-mapping\">What Is Domain Mapping?</h2>\n<p>Domain mapping goes beyond URL seeding. Instead of checking a single sitemap or index, <code>DomainMapper</code> combines <strong>8 discovery sources</strong> to find every URL under a domain — including subdomains you didn't know existed.</p>\n<h3 id=\"domainmapper-vs-asyncurlseeder\">DomainMapper vs AsyncUrlSeeder</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Aspect</th>\n<th>AsyncUrlSeeder</th>\n<th>DomainMapper</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Scope</strong></td>\n<td>Single host, listed URLs only</td>\n<td>Entire domain + all subdomains</td>\n</tr>\n<tr>\n<td><strong>Sources</strong></td>\n<td>Sitemap + Common Crawl</td>\n<td>8 sources (sitemap, CC, Wayback, crt.sh, probe, robots.txt, feeds, homepage)</td>\n</tr>\n<tr>\n<td><strong>Subdomain discovery</strong></td>\n<td>No</td>\n<td>Yes (Certificate Transparency, DNS, Wayback)</td>\n</tr>\n<tr>\n<td><strong>Soft-404 detection</strong></td>\n<td>No</td>\n<td>Yes (fingerprints SPA sites)</td>\n</tr>\n<tr>\n<td><strong>Best for</strong></td>\n<td>Known domains with good sitemaps</td>\n<td>Full domain reconnaissance</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Real-world example</strong>: For <code>superdesign.dev</code>, AsyncUrlSeeder found 4 URLs. DomainMapper found <strong>171 URLs across 11 hosts</strong> — including docs, API servers, staging environments, and analytics dashboards that no sitemap listed.</p>\n<h2 id=\"quick-start\">Quick Start</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> DomainMapper, DomainMapperConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> DomainMapper() <span class=\"hljs-keyword\">as</span> mapper:\n        results = <span class=\"hljs-keyword\">await</span> mapper.scan(<span class=\"hljs-string\">\"example.com\"</span>)\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> URLs\"</span>)\n    <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> results[:<span class=\"hljs-number\">10</span>]:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  [<span class=\"hljs-subst\">{r[<span class=\"hljs-string\">'source'</span>]}</span>] <span class=\"hljs-subst\">{r[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n        <span class=\"hljs-keyword\">if</span> r.get(<span class=\"hljs-string\">\"head_data\"</span>, {}).get(<span class=\"hljs-string\">\"title\"</span>):\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"    Title: <span class=\"hljs-subst\">{r[<span class=\"hljs-string\">'head_data'</span>][<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Or via <code>AsyncWebCrawler</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai import AsyncWebCrawler, <span class=\"hljs-function\">DomainMapperConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    results</span> = <span class=\"hljs-keyword\">await</span> crawler.amap_domain(<span class=\"hljs-string\">\"example.com\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"the-8-discovery-sources\">The 8 Discovery Sources</h2>\n<p>DomainMapper combines these sources, each catching URLs the others miss:</p>\n<h3 id=\"1-sitemap-sitemap-discovery\">1. <code>sitemap</code> — Sitemap Discovery</h3>\n<p>Checks <code>/sitemap.xml</code>, <code>/sitemap_index.xml</code>, and <code>robots.txt</code> <code>Sitemap:</code> directives <strong>on every discovered host</strong> — not just the root domain.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-cc-common-crawl\">2. <code>cc</code> — Common Crawl</h3>\n<p>Queries the Common Crawl CDX API for <code>*.domain.tld/*</code>, catching URLs and subdomains the web's largest public crawl has indexed.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"cc\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-wayback-wayback-machine\">3. <code>wayback</code> — Wayback Machine</h3>\n<p>Queries the Internet Archive's CDX API. Often has different coverage than Common Crawl — including historical pages that have since been removed.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"wayback\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"4-crt-certificate-transparency\">4. <code>crt</code> — Certificate Transparency</h3>\n<p>Queries <a href=\"https://crt.sh\">crt.sh</a> for SSL certificates issued to <code>*.domain.tld</code>. This is the single most effective subdomain discovery technique — it found 14 subdomains for <code>superdesign.dev</code> that no other source knew about.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"crt\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"5-probe-common-path-probing\">5. <code>probe</code> — Common Path Probing</h3>\n<p>Tries ~25 well-known paths on each discovered host (<code>/docs</code>, <code>/api</code>, <code>/login</code>, <code>/dashboard</code>, <code>/openapi.json</code>, etc.). Combined with soft-404 detection to avoid false positives.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">config = DomainMapperConfig(<span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"probe\"</span>)\n\n<span class=\"hljs-comment\"># Add custom paths to probe</span>\nconfig = DomainMapperConfig(\n    <span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"probe\"</span>,\n    probe_paths=[<span class=\"hljs-string\">\"/custom-api\"</span>, <span class=\"hljs-string\">\"/internal/status\"</span>]\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"6-robots-robotstxt-path-mining\">6. <code>robots</code> — robots.txt Path Mining</h3>\n<p>Parses <code>Disallow:</code> and <code>Allow:</code> lines from <code>robots.txt</code>. These are confirmed real paths the site acknowledges exist — often revealing admin panels, APIs, and internal tools that aren't linked anywhere.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"robots\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"7-feed-rssatom-feed-parsing\">7. <code>feed</code> — RSS/Atom Feed Parsing</h3>\n<p>Discovers and parses RSS/Atom feeds at common paths (<code>/feed</code>, <code>/rss</code>, <code>/atom.xml</code>, etc.). Feeds are curated lists of content URLs maintained by the site.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"feed\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"8-homepage-homepage-link-extraction\">8. <code>homepage</code> — Homepage Link Extraction</h3>\n<p>Fetches each host's homepage via HTTP and extracts all internal links using <code>quick_extract_links()</code>. Also mines <code>&lt;link rel=\"alternate|preload|prefetch\"&gt;</code> tags from the <code>&lt;head&gt;</code> for additional URLs. No browser needed.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">config</span> = DomainMapperConfig(source=<span class=\"hljs-string\">\"homepage\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"combining-sources\">Combining Sources</h3>\n<p>Sources are combined with <code>+</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Default: most useful combination</span>\nconfig = DomainMapperConfig(<span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"sitemap+cc+crt+probe\"</span>)\n\n<span class=\"hljs-comment\"># Maximum coverage: all 8 sources</span>\nconfig = DomainMapperConfig(\n    <span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"sitemap+cc+wayback+crt+probe+robots+feed+homepage\"</span>\n)\n\n<span class=\"hljs-comment\"># Lightweight: just sitemap + probing</span>\nconfig = DomainMapperConfig(<span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"sitemap+probe\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"how-it-works-the-three-phases\">How It Works: The Three Phases</h2>\n<h3 id=\"phase-1-host-discovery\">Phase 1: Host Discovery</h3>\n<p>DomainMapper first discovers all subdomains under your domain:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">superdesign.dev\n├── crt.sh           → docs, app, cloud, insights, staging-api, ui2web, ...\n├── Wayback CDX      → api, app, docs, www, ...\n├── Common Crawl     → app, www, ...\n└── DNS guessing     → www, app, api, docs, blog, admin, cloud, ...\n\n<span class=\"hljs-section\">Result: 13 validated hosts</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Each discovered host is validated with an HTTP HEAD request. Hosts that don't respond are dropped.</p>\n<h3 id=\"phase-2-per-host-scanning\">Phase 2: Per-Host Scanning</h3>\n<p>For each validated host, DomainMapper runs all enabled sources in parallel:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">docs.superdesign.dev\n├── Soft-<span class=\"hljs-number\">404</span> fingerprint  → (<span class=\"hljs-number\">404</span> returns proper error — <span class=\"hljs-keyword\">no</span> SPA issue)\n├── robots.txt            → <span class=\"hljs-number\">1</span> sitemap URL, <span class=\"hljs-number\">1</span> disallow path\n├── Sitemap parsing       → <span class=\"hljs-number\">19</span> URLs\n├── Path probing          → <span class=\"hljs-number\">2</span> valid (<span class=\"hljs-regexp\">/docs, /</span>)\n├── Feed discovery        → (<span class=\"hljs-keyword\">no</span> feeds found)\n└── Homepage extraction   → <span class=\"hljs-number\">26</span> internal links\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"phase-3-post-processing\">Phase 3: Post-Processing</h3>\n<p>All discovered URLs go through:</p>\n<ol>\n<li><strong>URL normalization</strong> — using <code>normalize_url()</code> to canonicalize</li>\n<li><strong>Deduplication</strong> — by normalized URL, merging source attribution</li>\n<li><strong>Nonsense filtering</strong> — removes static assets (JS, CSS, images, fonts), webpack chunks, Wayback garbage</li>\n<li><strong>Head extraction</strong> — parallel <code>&lt;head&gt;</code> fetching for metadata (optional)</li>\n<li><strong>BM25 scoring</strong> — relevance scoring against a query (optional)</li>\n</ol>\n<h2 id=\"soft-404-detection\">Soft-404 Detection</h2>\n<p>Many modern SPAs return HTTP 200 for every URL — even pages that don't exist. DomainMapper detects this:</p>\n<ol>\n<li><strong>Fingerprinting</strong>: Fetches a guaranteed-nonexistent URL (e.g., <code>/c4ai-probe-a1b2c3d4</code>) on each host</li>\n<li><strong>Recording</strong>: Captures the response title and body hash</li>\n<li><strong>Filtering</strong>: When probing real paths, compares against the fingerprint. If they match → soft-404, filtered out</li>\n</ol>\n<p>For <code>superdesign.dev</code>, this correctly:\n- Blocked <strong>all 25+ probe paths</strong> on <code>app.superdesign.dev</code> (SPA that returns 200 for everything)\n- Blocked <strong>476 sitemap URLs</strong> from <code>app.superdesign.dev</code> (all rendering the same shell)\n- Kept all 19 legitimate URLs from <code>docs.superdesign.dev</code></p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Soft-404 detection is on by default</span>\n<span class=\"hljs-attr\">config</span> = DomainMapperConfig(soft_404_detection=<span class=\"hljs-literal\">True</span>)\n\n<span class=\"hljs-comment\"># Disable if you want raw results</span>\n<span class=\"hljs-attr\">config</span> = DomainMapperConfig(soft_404_detection=<span class=\"hljs-literal\">False</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"configuration-reference\">Configuration Reference</h2>\n<h3 id=\"domainmapperconfig\">DomainMapperConfig</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Type</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>source</code></td>\n<td>str</td>\n<td><code>\"sitemap+cc+crt+probe\"</code></td>\n<td>Discovery sources joined by <code>+</code></td>\n</tr>\n<tr>\n<td><code>max_urls</code></td>\n<td>int</td>\n<td><code>-1</code></td>\n<td>Maximum URLs to return (-1 = unlimited)</td>\n</tr>\n<tr>\n<td><code>concurrency</code></td>\n<td>int</td>\n<td><code>50</code></td>\n<td>Max concurrent requests across all hosts</td>\n</tr>\n<tr>\n<td><code>hits_per_sec</code></td>\n<td>int</td>\n<td><code>10</code></td>\n<td>Rate limit in requests/second</td>\n</tr>\n<tr>\n<td><code>force</code></td>\n<td>bool</td>\n<td><code>False</code></td>\n<td>Bypass all caches</td>\n</tr>\n<tr>\n<td><code>extract_head</code></td>\n<td>bool</td>\n<td><code>True</code></td>\n<td>Fetch and parse <code>&lt;head&gt;</code> metadata</td>\n</tr>\n<tr>\n<td><code>filter_nonsense_urls</code></td>\n<td>bool</td>\n<td><code>True</code></td>\n<td>Filter static assets and utility URLs</td>\n</tr>\n<tr>\n<td><code>soft_404_detection</code></td>\n<td>bool</td>\n<td><code>True</code></td>\n<td>Fingerprint and filter soft-404 pages</td>\n</tr>\n<tr>\n<td><code>query</code></td>\n<td>str</td>\n<td><code>None</code></td>\n<td>BM25 relevance query (requires <code>extract_head=True</code>)</td>\n</tr>\n<tr>\n<td><code>score_threshold</code></td>\n<td>float</td>\n<td><code>None</code></td>\n<td>Minimum relevance score (0.0-1.0)</td>\n</tr>\n<tr>\n<td><code>scoring_method</code></td>\n<td>str</td>\n<td><code>\"bm25\"</code></td>\n<td>Scoring algorithm</td>\n</tr>\n<tr>\n<td><code>probe_paths</code></td>\n<td>List[str]</td>\n<td><code>None</code></td>\n<td>Extra paths to probe on each host</td>\n</tr>\n<tr>\n<td><code>common_subdomains</code></td>\n<td>List[str]</td>\n<td><code>None</code></td>\n<td>Extra subdomain prefixes to guess</td>\n</tr>\n<tr>\n<td><code>use_browser_for_homepage</code></td>\n<td>bool</td>\n<td><code>False</code></td>\n<td>Use Playwright for JS-rendered homepages</td>\n</tr>\n<tr>\n<td><code>verbose</code></td>\n<td>bool</td>\n<td><code>None</code></td>\n<td>Override logger verbose setting</td>\n</tr>\n<tr>\n<td><code>cache_ttl_hours</code></td>\n<td>int</td>\n<td><code>24</code></td>\n<td>Hours before cached results expire</td>\n</tr>\n<tr>\n<td><code>dns_timeout</code></td>\n<td>float</td>\n<td><code>3.0</code></td>\n<td>Timeout for DNS resolution (seconds)</td>\n</tr>\n<tr>\n<td><code>http_timeout</code></td>\n<td>float</td>\n<td><code>10.0</code></td>\n<td>Timeout for HTTP requests (seconds)</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"output-format\">Output Format</h3>\n<p>Each result is a dict:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">{\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://docs.superdesign.dev/quickstart\"</span>,\n    <span class=\"hljs-string\">\"host\"</span>: <span class=\"hljs-string\">\"docs.superdesign.dev\"</span>,\n    <span class=\"hljs-string\">\"source\"</span>: <span class=\"hljs-string\">\"homepage+sitemap\"</span>,     <span class=\"hljs-comment\"># which source(s) found it</span>\n    <span class=\"hljs-string\">\"status\"</span>: <span class=\"hljs-string\">\"valid\"</span>,                <span class=\"hljs-comment\"># valid | not_valid | soft_404</span>\n    <span class=\"hljs-string\">\"head_data\"</span>: {                    <span class=\"hljs-comment\"># if extract_head=True</span>\n        <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"Quickstart\"</span>,\n        <span class=\"hljs-string\">\"meta\"</span>: {<span class=\"hljs-string\">\"description\"</span>: <span class=\"hljs-string\">\"...\"</span>},\n        <span class=\"hljs-string\">\"link\"</span>: {...},\n        <span class=\"hljs-string\">\"jsonld\"</span>: [...]\n    },\n    <span class=\"hljs-string\">\"relevance_score\"</span>: 0.85,          <span class=\"hljs-comment\"># if query provided</span>\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"practical-examples\">Practical Examples</h2>\n<h3 id=\"discover-and-crawl-documentation\">Discover and Crawl Documentation</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, DomainMapperConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_all_docs</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Step 1: Discover all URLs</span>\n        pages = <span class=\"hljs-keyword\">await</span> crawler.amap_domain(<span class=\"hljs-string\">\"example.com\"</span>, DomainMapperConfig(\n            source=<span class=\"hljs-string\">\"sitemap+crt+probe+homepage\"</span>,\n            extract_head=<span class=\"hljs-literal\">True</span>,\n            query=<span class=\"hljs-string\">\"documentation tutorial guide\"</span>,\n        ))\n\n        <span class=\"hljs-comment\"># Step 2: Filter for docs</span>\n        doc_urls = [\n            p[<span class=\"hljs-string\">\"url\"</span>] <span class=\"hljs-keyword\">for</span> p <span class=\"hljs-keyword\">in</span> pages\n            <span class=\"hljs-keyword\">if</span> p.get(<span class=\"hljs-string\">\"relevance_score\"</span>, <span class=\"hljs-number\">0</span>) &gt; <span class=\"hljs-number\">0.3</span>\n        ]\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(doc_urls)}</span> documentation pages\"</span>)\n\n        <span class=\"hljs-comment\"># Step 3: Crawl them</span>\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(\n            doc_urls[:<span class=\"hljs-number\">50</span>],\n            config=CrawlerRunConfig(only_text=<span class=\"hljs-literal\">True</span>)\n        )\n        <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-keyword\">if</span> r.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Crawled: <span class=\"hljs-subst\">{r.url}</span>\"</span>)\n\nasyncio.run(crawl_all_docs())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"security-audit-find-exposed-services\">Security Audit: Find Exposed Services</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">audit_domain</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> DomainMapper() <span class=\"hljs-keyword\">as</span> mapper:\n        results = <span class=\"hljs-keyword\">await</span> mapper.scan(<span class=\"hljs-string\">\"company.com\"</span>, DomainMapperConfig(\n            source=<span class=\"hljs-string\">\"crt+probe+robots\"</span>,\n            extract_head=<span class=\"hljs-literal\">True</span>,\n            probe_paths=[\n                <span class=\"hljs-string\">\"/openapi.json\"</span>, <span class=\"hljs-string\">\"/swagger.json\"</span>, <span class=\"hljs-string\">\"/api-docs\"</span>,\n                <span class=\"hljs-string\">\"/graphql\"</span>, <span class=\"hljs-string\">\"/.env\"</span>, <span class=\"hljs-string\">\"/debug\"</span>, <span class=\"hljs-string\">\"/admin\"</span>,\n                <span class=\"hljs-string\">\"/phpinfo.php\"</span>, <span class=\"hljs-string\">\"/server-status\"</span>,\n            ],\n        ))\n\n        <span class=\"hljs-comment\"># Flag exposed services</span>\n        <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> results:\n            title = r.get(<span class=\"hljs-string\">\"head_data\"</span>, {}).get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"\"</span>)\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">any</span>(x <span class=\"hljs-keyword\">in</span> title.lower() <span class=\"hljs-keyword\">for</span> x <span class=\"hljs-keyword\">in</span> [<span class=\"hljs-string\">\"swagger\"</span>, <span class=\"hljs-string\">\"api\"</span>, <span class=\"hljs-string\">\"admin\"</span>, <span class=\"hljs-string\">\"debug\"</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  EXPOSED: <span class=\"hljs-subst\">{r[<span class=\"hljs-string\">'url'</span>]}</span> — <span class=\"hljs-subst\">{title}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"compare-subdomains-across-a-domain\">Compare Subdomains Across a Domain</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">map_infrastructure</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> DomainMapper() <span class=\"hljs-keyword\">as</span> mapper:\n        results = <span class=\"hljs-keyword\">await</span> mapper.scan(<span class=\"hljs-string\">\"company.com\"</span>, DomainMapperConfig(\n            source=<span class=\"hljs-string\">\"crt+probe\"</span>,\n            extract_head=<span class=\"hljs-literal\">False</span>,\n        ))\n\n        <span class=\"hljs-comment\"># Group by host</span>\n        <span class=\"hljs-keyword\">from</span> collections <span class=\"hljs-keyword\">import</span> defaultdict\n        by_host = defaultdict(<span class=\"hljs-built_in\">list</span>)\n        <span class=\"hljs-keyword\">for</span> r <span class=\"hljs-keyword\">in</span> results:\n            by_host[r[<span class=\"hljs-string\">\"host\"</span>]].append(r)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Discovered <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(by_host)}</span> hosts:\"</span>)\n        <span class=\"hljs-keyword\">for</span> host, urls <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">sorted</span>(by_host.items()):\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  <span class=\"hljs-subst\">{host}</span>: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(urls)}</span> URLs\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"tips-and-best-practices\">Tips and Best Practices</h2>\n<ol>\n<li>\n<p><strong>Start with the default sources</strong> (<code>sitemap+cc+crt+probe</code>). Add <code>wayback</code>, <code>robots</code>, <code>feed</code>, and <code>homepage</code> if you need maximum coverage.</p>\n</li>\n<li>\n<p><strong>Use <code>extract_head=False</code> for speed</strong> when you just need URL lists. Head extraction makes ~1 HTTP request per URL.</p>\n</li>\n<li>\n<p><strong>The <code>query</code> parameter is powerful</strong> for finding specific content across a large domain without crawling anything.</p>\n</li>\n<li>\n<p><strong><code>probe_paths</code> is your extensibility hook</strong> — add domain-specific paths you suspect exist.</p>\n</li>\n<li>\n<p><strong>Rate limiting matters</strong> — <code>hits_per_sec=10</code> is respectful. Lower it for smaller sites, raise it for your own infrastructure.</p>\n</li>\n<li>\n<p><strong>Soft-404 detection is critical for SPAs</strong> — without it, single-page apps flood your results with hundreds of identical shell pages.</p>\n</li>\n</ol>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"../url-seeding/\">URL Seeding</a> — simpler, single-host URL discovery from sitemaps and Common Crawl</li>\n<li><a href=\"../deep-crawling/\">Deep Crawling</a> — follow links dynamically within pages</li>\n<li><a href=\"../../advanced/multi-url-crawling/\">Multi-URL Crawling</a> — crawl discovered URLs in bulk</li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/examples/",
    "title": "Code Examples - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"code-examples\">Code Examples</h1>\n<p>This page provides a comprehensive list of example scripts that demonstrate various features and capabilities of Crawl4AI. Each example is designed to showcase specific functionality, making it easier for you to understand how to implement these features in your own projects.</p>\n<h2 id=\"getting-started-examples\">Getting Started Examples</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Hello World</td>\n<td>A simple introductory example demonstrating basic usage of AsyncWebCrawler with JavaScript execution and content filtering.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/hello_world.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Quickstart</td>\n<td>A comprehensive collection of examples showcasing various features including basic crawling, content cleaning, link analysis, JavaScript execution, CSS selectors, media handling, custom hooks, proxy configuration, screenshots, and multiple extraction strategies.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/quickstart.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Quickstart Set 1</td>\n<td>Basic examples for getting started with Crawl4AI.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/quickstart_examples_set_1.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Quickstart Set 2</td>\n<td>More advanced examples for working with Crawl4AI.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/quickstart_examples_set_2.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"proxies\">Proxies</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>NSTProxy</strong></td>\n<td><a href=\"https://www.nstproxy.com/?utm_source=crawl4ai\">NSTProxy</a> Seamlessly integrates with crawl4ai — no setup required. Access high-performance residential, datacenter, ISP, and IPv6 proxies with smart rotation and anti-blocking technology. Starts from $0.1/GB. Use code crawl4ai for 10% off.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/tree/main/docs/examples/proxy\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"browser-crawling-features\">Browser &amp; Crawling Features</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Built-in Browser</td>\n<td>Demonstrates how to use the built-in browser capabilities.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/builtin_browser_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Browser Optimization</td>\n<td>Focuses on browser performance optimization techniques.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/browser_optimization_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>arun vs arun_many</td>\n<td>Compares the <code>arun</code> and <code>arun_many</code> methods for single vs. multiple URL crawling.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/arun_vs_arun_many.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Multiple URLs</td>\n<td>Shows how to crawl multiple URLs asynchronously.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/async_webcrawler_multiple_urls_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Page Interaction</td>\n<td>Guide on interacting with dynamic elements through clicks.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/tutorial_dynamic_clicks.md\">View Guide</a></td>\n</tr>\n<tr>\n<td>Crawler Monitor</td>\n<td>Shows how to monitor the crawler's activities and status.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/crawler_monitor_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Full Page Screenshot &amp; PDF</td>\n<td>Guide on capturing full-page screenshots and PDFs from massive webpages.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/full_page_screenshot_and_pdf_export.md\">View Guide</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"advanced-crawling-deep-crawling\">Advanced Crawling &amp; Deep Crawling</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Deep Crawling</td>\n<td>An extensive tutorial on deep crawling capabilities, demonstrating BFS and BestFirst strategies, stream vs. non-stream execution, filters, scorers, and advanced configurations.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/deepcrawl_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Virtual Scroll</td>\n<td>Comprehensive examples for handling virtualized scrolling on sites like Twitter, Instagram. Demonstrates different scrolling scenarios with local test server.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/virtual_scroll_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Adaptive Crawling</td>\n<td>Demonstrates intelligent crawling that automatically determines when sufficient information has been gathered.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/adaptive_crawling/\">View Code</a></td>\n</tr>\n<tr>\n<td>Dispatcher</td>\n<td>Shows how to use the crawl dispatcher for advanced workload management.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/dispatcher_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Storage State</td>\n<td>Tutorial on managing browser storage state for persistence.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/storage_state_tutorial.md\">View Guide</a></td>\n</tr>\n<tr>\n<td>Network Console Capture</td>\n<td>Demonstrates how to capture and analyze network requests and console logs.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/network_console_capture_example.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"extraction-strategies\">Extraction Strategies</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Extraction Strategies</td>\n<td>Demonstrates different extraction strategies with various input formats (markdown, HTML, fit_markdown) and JSON-based extractors (CSS and XPath).</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/extraction_strategies_examples.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Scraping Strategies</td>\n<td>Compares the performance of different scraping strategies.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/scraping_strategies_performance.py\">View Code</a></td>\n</tr>\n<tr>\n<td>LLM Extraction</td>\n<td>Demonstrates LLM-based extraction specifically for OpenAI pricing data.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/llm_extraction_openai_pricing.py\">View Code</a></td>\n</tr>\n<tr>\n<td>LLM Markdown</td>\n<td>Shows how to use LLMs to generate markdown from crawled content.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/llm_markdown_generator.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Summarize Page</td>\n<td>Shows how to summarize web page content.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/summarize_page.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"e-commerce-specialized-crawling\">E-commerce &amp; Specialized Crawling</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Amazon Product Extraction</td>\n<td>Demonstrates how to extract structured product data from Amazon search results using CSS selectors.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/amazon_product_extraction_direct_url.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Amazon with Hooks</td>\n<td>Shows how to use hooks with Amazon product extraction.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/amazon_product_extraction_using_hooks.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Amazon with JavaScript</td>\n<td>Demonstrates using custom JavaScript for Amazon product extraction.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/amazon_product_extraction_using_use_javascript.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Crypto Analysis</td>\n<td>Demonstrates how to crawl and analyze cryptocurrency data.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/crypto_analysis_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>SERP API</td>\n<td>Demonstrates using Crawl4AI with search engine result pages.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/serp_api_project_11_feb.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"anti-bot-stealth-features\">Anti-Bot &amp; Stealth Features</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Stealth Mode Quick Start</td>\n<td>Five practical examples showing how to use stealth mode for bypassing basic bot detection.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/stealth_mode_quick_start.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Stealth Mode Comprehensive</td>\n<td>Comprehensive demonstration of stealth mode features with bot detection testing and comparisons.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/stealth_mode_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Undetected Browser</td>\n<td>Simple example showing how to use the undetected browser adapter.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/hello_world_undetected.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Undetected Browser Demo</td>\n<td>Basic demo comparing regular and undetected browser modes.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/undetected_simple_demo.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Undetected Tests</td>\n<td>Advanced tests comparing regular vs undetected browsers on various bot detection services.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/tree/main/docs/examples/undetectability/\">View Folder</a></td>\n</tr>\n<tr>\n<td>CapSolver Captcha Solver</td>\n<td>Seamlessly integrate with <a href=\"https://www.capsolver.com/?utm_source=crawl4ai&amp;utm_medium=github_pr&amp;utm_campaign=crawl4ai_integration\">CapSolver</a> to automatically solve reCAPTCHA v2/v3, Cloudflare Turnstile / Challenges, AWS WAF and more for uninterrupted scraping and automation.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/tree/main/docs/examples/capsolver_captcha_solver/\">View Folder</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"customization-security\">Customization &amp; Security</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Hooks</td>\n<td>Illustrates how to use hooks at different stages of the crawling process for advanced customization.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/hooks_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Identity-Based Browsing</td>\n<td>Illustrates identity-based browsing configurations for authentic browsing experiences.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/identity_based_browsing.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Proxy Rotation</td>\n<td>Shows how to use proxy rotation for web scraping and avoiding IP blocks.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/proxy_rotation_demo.py\">View Code</a></td>\n</tr>\n<tr>\n<td>SSL Certificate</td>\n<td>Illustrates SSL certificate handling and verification.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/ssl_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Language Support</td>\n<td>Shows how to handle different languages during crawling.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/language_support_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Geolocation</td>\n<td>Demonstrates how to use geolocation features.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/use_geo_location.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"docker-deployment\">Docker &amp; Deployment</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Docker Config</td>\n<td>Demonstrates how to create and use Docker configuration objects.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/docker_config_obj.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Docker Basic</td>\n<td>A test suite for Docker deployment, showcasing various functionalities through the Docker API.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/docker_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Docker REST API</td>\n<td>Shows how to interact with Crawl4AI Docker using REST API calls.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/docker_python_rest_api.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Docker SDK</td>\n<td>Demonstrates using the Python SDK for Crawl4AI Docker.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/docker_python_sdk.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"application-examples\">Application Examples</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Research Assistant</td>\n<td>Demonstrates how to build a research assistant using Crawl4AI.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/research_assistant.py\">View Code</a></td>\n</tr>\n<tr>\n<td>REST Call</td>\n<td>Shows how to make REST API calls with Crawl4AI.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/rest_call.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Chainlit Integration</td>\n<td>Shows how to integrate Crawl4AI with Chainlit.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/chainlit.md\">View Guide</a></td>\n</tr>\n<tr>\n<td>Crawl4AI vs FireCrawl</td>\n<td>Compares Crawl4AI with the FireCrawl library.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/crawlai_vs_firecrawl.py\">View Code</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"content-generation-markdown\">Content Generation &amp; Markdown</h2>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Example</th>\n<th>Description</th>\n<th>Link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Content Source</td>\n<td>Demonstrates how to work with different content sources in markdown generation.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/markdown/content_source_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Content Source (Short)</td>\n<td>A simplified version of content source usage.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/markdown/content_source_short_example.py\">View Code</a></td>\n</tr>\n<tr>\n<td>Built-in Browser Guide</td>\n<td>Guide for using the built-in browser capabilities.</td>\n<td><a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/README_BUILTIN_BROWSER.md\">View Guide</a></td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"running-the-examples\">Running the Examples</h2>\n<p>To run any of these examples, you'll need to have Crawl4AI installed:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">pip install crawl4ai\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Then, you can run an example script like this:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">python -m docs.examples.hello_world\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>For examples that require additional dependencies or environment variables, refer to the comments at the top of each file.</p>\n<p>Some examples may require:\n- API keys (for LLM-based examples)\n- Docker setup (for Docker-related examples)\n- Additional dependencies (specified in the example files)</p>\n<h2 id=\"contributing-new-examples\">Contributing New Examples</h2>\n<p>If you've created an interesting example that demonstrates a unique use case or feature of Crawl4AI, we encourage you to contribute it to our examples collection. Please see our <a href=\"https://github.com/unclecode/crawl4ai/blob/main/CONTRIBUTORS.md\">contribution guidelines</a> for more information.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/fit-markdown/",
    "title": "Fit Markdown - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"fit-markdown-with-pruning-bm25\">Fit Markdown with Pruning &amp; BM25</h1>\n<p><strong>Fit Markdown</strong> is a specialized <strong>filtered</strong> version of your page’s markdown, focusing on the most relevant content. By default, Crawl4AI converts the entire HTML into a broad <strong>raw_markdown</strong>. With fit markdown, we apply a <strong>content filter</strong> algorithm (e.g., <strong>Pruning</strong> or <strong>BM25</strong>) to remove or rank low-value sections—such as repetitive sidebars, shallow text blocks, or irrelevancies—leaving a concise textual “core.”</p>\n<hr>\n<h2 id=\"1-how-fit-markdown-works\">1. How “Fit Markdown” Works</h2>\n<h3 id=\"11-the-content_filter\">1.1 The <code>content_filter</code></h3>\n<p>In <strong><code>CrawlerRunConfig</code></strong>, you can specify a <strong><code>content_filter</code></strong> to shape how content is pruned or ranked before final markdown generation. A filter’s logic is applied <strong>before</strong> or <strong>during</strong> the HTML→Markdown process, producing:</p>\n<ul>\n<li><strong><code>result.markdown.raw_markdown</code></strong> (unfiltered)</li>\n<li><strong><code>result.markdown.fit_markdown</code></strong> (filtered or “fit” version)</li>\n<li><strong><code>result.markdown.fit_html</code></strong> (the corresponding HTML snippet that produced <code>fit_markdown</code>)</li>\n</ul>\n<h3 id=\"12-common-filters\">1.2 Common Filters</h3>\n<p>1. <strong>PruningContentFilter</strong> – Scores each node by text density, link density, and tag importance, discarding those below a threshold.<br>\n2. <strong>BM25ContentFilter</strong> – Focuses on textual relevance using BM25 ranking, especially useful if you have a specific user query (e.g., “machine learning” or “food nutrition”).</p>\n<hr>\n<h2 id=\"2-pruningcontentfilter\">2. PruningContentFilter</h2>\n<p><strong>Pruning</strong> discards less relevant nodes based on <strong>text density, link density, and tag importance</strong>. It’s a heuristic-based approach—if certain sections appear too “thin” or too “spammy,” they’re pruned.</p>\n<h3 id=\"21-usage-example\">2.1 Usage Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Step 1: Create a pruning filter</span>\n    prune_filter = PruningContentFilter(\n        <span class=\"hljs-comment\"># Lower → more content retained, higher → more content pruned</span>\n        threshold=<span class=\"hljs-number\">0.45</span>,           \n        <span class=\"hljs-comment\"># \"fixed\" or \"dynamic\"</span>\n        threshold_type=<span class=\"hljs-string\">\"dynamic\"</span>,  \n        <span class=\"hljs-comment\"># Ignore nodes with &lt;5 words</span>\n        min_word_threshold=<span class=\"hljs-number\">5</span>      \n    )\n\n    <span class=\"hljs-comment\"># Step 2: Insert it into a Markdown Generator</span>\n    md_generator = DefaultMarkdownGenerator(content_filter=prune_filter)\n\n    <span class=\"hljs-comment\"># Step 3: Pass it to CrawlerRunConfig</span>\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, \n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># 'fit_markdown' is your pruned content, focusing on \"denser\" text</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.fit_markdown))\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"22-key-parameters\">2.2 Key Parameters</h3>\n<ul>\n<li><strong><code>min_word_threshold</code></strong> (int): If a block has fewer words than this, it’s pruned.  </li>\n<li><strong><code>threshold_type</code></strong> (str):</li>\n<li><code>\"fixed\"</code> → each node must exceed <code>threshold</code> (0–1).  </li>\n<li><code>\"dynamic\"</code> → node scoring adjusts according to tag type, text/link density, etc.  </li>\n<li><strong><code>threshold</code></strong> (float, default ~0.48): The base or “anchor” cutoff.  </li>\n</ul>\n<p><strong>Algorithmic Factors</strong>:</p>\n<ul>\n<li><strong>Text density</strong> – Encourages blocks that have a higher ratio of text to overall content.  </li>\n<li><strong>Link density</strong> – Penalizes sections that are mostly links.  </li>\n<li><strong>Tag importance</strong> – e.g., an <code>&lt;article&gt;</code> or <code>&lt;p&gt;</code> might be more important than a <code>&lt;div&gt;</code>.  </li>\n<li><strong>Structural context</strong> – If a node is deeply nested or in a suspected sidebar, it might be deprioritized.</li>\n</ul>\n<hr>\n<h2 id=\"3-bm25contentfilter\">3. BM25ContentFilter</h2>\n<p><strong>BM25</strong> is a classical text ranking algorithm often used in search engines. If you have a <strong>user query</strong> or rely on page metadata to derive a query, BM25 can identify which text chunks best match that query.</p>\n<h3 id=\"31-usage-example\">3.1 Usage Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1) A BM25 filter with a user query</span>\n    bm25_filter = BM25ContentFilter(\n        user_query=<span class=\"hljs-string\">\"startup fundraising tips\"</span>,\n        <span class=\"hljs-comment\"># Adjust for stricter or looser results</span>\n        bm25_threshold=<span class=\"hljs-number\">1.2</span>  \n    )\n\n    <span class=\"hljs-comment\"># 2) Insert into a Markdown Generator</span>\n    md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)\n\n    <span class=\"hljs-comment\"># 3) Pass to crawler config</span>\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, \n            config=config\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown (BM25 query-based):\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-parameters\">3.2 Parameters</h3>\n<ul>\n<li><strong><code>user_query</code></strong> (str, optional): E.g. <code>\"machine learning\"</code>. If blank, the filter tries to glean a query from page metadata.  </li>\n<li><strong><code>bm25_threshold</code></strong> (float, default 1.0):  </li>\n<li>Higher → fewer chunks but more relevant.  </li>\n<li>Lower → more inclusive.  </li>\n</ul>\n<blockquote>\n<p>In more advanced scenarios, you might see parameters like <code>language</code>, <code>case_sensitive</code>, or <code>priority_tags</code> to refine how text is tokenized or weighted.</p>\n</blockquote>\n<hr>\n<h2 id=\"4-accessing-the-fit-output\">4. Accessing the “Fit” Output</h2>\n<p>After the crawl, your “fit” content is found in <strong><code>result.markdown.fit_markdown</code></strong>. </p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">fit_md</span> = result.markdown.fit_markdown\n<span class=\"hljs-attr\">fit_html</span> = result.markdown.fit_html\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>If the content filter is <strong>BM25</strong>, you might see additional logic or references in <code>fit_markdown</code> that highlight relevant segments. If it’s <strong>Pruning</strong>, the text is typically well-cleaned but not necessarily matched to a query.</p>\n<hr>\n<h2 id=\"5-code-patterns-recap\">5. Code Patterns Recap</h2>\n<h3 id=\"51-pruning\">5.1 Pruning</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">prune_filter = PruningContentFilter(\n    threshold=0.5,\n    threshold_type=<span class=\"hljs-string\">\"fixed\"</span>,\n    min_word_threshold=10\n)\nmd_generator = DefaultMarkdownGenerator(content_filter=prune_filter)\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"52-bm25\">5.2 BM25</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">bm25_filter = BM25ContentFilter(\n    user_query=<span class=\"hljs-string\">\"health benefits fruit\"</span>,\n    bm25_threshold=1.2\n)\nmd_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"6-combining-with-word_count_threshold-exclusions\">6. Combining with “word_count_threshold” &amp; Exclusions</h2>\n<p>Remember you can also specify:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">10</span>,\n    excluded_tags<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"footer\"</span>, <span class=\"hljs-string\">\"header\"</span><span class=\"hljs-punctuation\">]</span>,\n    exclude_external_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    markdown_generator<span class=\"hljs-punctuation\">=</span>DefaultMarkdownGenerator<span class=\"hljs-punctuation\">(</span>\n        content_filter<span class=\"hljs-punctuation\">=</span>PruningContentFilter<span class=\"hljs-punctuation\">(</span>threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.5</span><span class=\"hljs-punctuation\">)</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Thus, <strong>multi-level</strong> filtering occurs:</p>\n<ol>\n<li>The crawler’s <code>excluded_tags</code> are removed from the HTML first.  </li>\n<li>The content filter (Pruning, BM25, or custom) prunes or ranks the remaining text blocks.  </li>\n<li>The final “fit” content is generated in <code>result.markdown.fit_markdown</code>.</li>\n</ol>\n<hr>\n<h2 id=\"7-custom-filters\">7. Custom Filters</h2>\n<p>If you need a different approach (like a specialized ML model or site-specific heuristics), you can create a new class inheriting from <code>RelevantContentFilter</code> and implement <code>filter_content(html)</code>. Then inject it into your <strong>markdown generator</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> RelevantContentFilter\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">MyCustomFilter</span>(<span class=\"hljs-title class_ inherited__\">RelevantContentFilter</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">filter_content</span>(<span class=\"hljs-params\">self, html, min_word_threshold=<span class=\"hljs-literal\">None</span></span>):\n        <span class=\"hljs-comment\"># parse HTML, implement custom logic</span>\n        <span class=\"hljs-keyword\">return</span> [block <span class=\"hljs-keyword\">for</span> block <span class=\"hljs-keyword\">in</span> ... <span class=\"hljs-keyword\">if</span> ... some condition...]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Steps</strong>:</p>\n<ol>\n<li>Subclass <code>RelevantContentFilter</code>.  </li>\n<li>Implement <code>filter_content(...)</code>.  </li>\n<li>Use it in your <code>DefaultMarkdownGenerator(content_filter=MyCustomFilter(...))</code>.</li>\n</ol>\n<hr>\n<h2 id=\"8-final-thoughts\">8. Final Thoughts</h2>\n<p><strong>Fit Markdown</strong> is a crucial feature for:</p>\n<ul>\n<li><strong>Summaries</strong>: Quickly get the important text from a cluttered page.  </li>\n<li><strong>Search</strong>: Combine with <strong>BM25</strong> to produce content relevant to a query.  </li>\n<li><strong>AI Pipelines</strong>: Filter out boilerplate so LLM-based extraction or summarization runs on denser text.</li>\n</ul>\n<p><strong>Key Points</strong>:\n- <strong>PruningContentFilter</strong>: Great if you just want the “meatiest” text without a user query.<br>\n- <strong>BM25ContentFilter</strong>: Perfect for query-based extraction or searching.<br>\n- Combine with <strong><code>excluded_tags</code>, <code>exclude_external_links</code>, <code>word_count_threshold</code></strong> to refine your final “fit” text.<br>\n- Fit markdown ends up in <strong><code>result.markdown.fit_markdown</code></strong>; eventually <strong><code>result.markdown.fit_markdown</code></strong> in future versions.</p>\n<p>With these tools, you can <strong>zero in</strong> on the text that truly matters, ignoring spammy or boilerplate content, and produce a concise, relevant “fit markdown” for your AI or data pipelines. Happy pruning and searching!</p>\n<ul>\n<li>Last Updated: 2025-01-01</li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/installation/",
    "title": "Installation - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"installation-setup-2023-edition\">Installation &amp; Setup (2023 Edition)</h1>\n<h2 id=\"1-basic-installation\">1. Basic Installation</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">pip install crawl4ai\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This installs the <strong>core</strong> Crawl4AI library along with essential dependencies. <strong>No</strong> advanced features (like transformers or PyTorch) are included yet.</p>\n<h2 id=\"2-initial-setup-diagnostics\">2. Initial Setup &amp; Diagnostics</h2>\n<h3 id=\"21-run-the-setup-command\">2.1 Run the Setup Command</h3>\n<p>After installing, call:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What does it do?</strong>\n- Installs or updates required browser dependencies for both regular and undetected modes\n- Performs OS-level checks (e.g., missing libs on Linux)\n- Confirms your environment is ready to crawl</p>\n<h3 id=\"22-diagnostics\">2.2 Diagnostics</h3>\n<p>Optionally, you can run <strong>diagnostics</strong> to confirm everything is functioning:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-doctor\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This command attempts to:\n- Check Python version compatibility\n- Verify Playwright installation\n- Inspect environment variables or library conflicts</p>\n<p>If any issues arise, follow its suggestions (e.g., installing additional system packages) and re-run <code>crawl4ai-setup</code>.</p>\n<hr>\n<h2 id=\"3-verifying-installation-a-simple-crawl-skip-this-step-if-you-already-run-crawl4ai-doctor\">3. Verifying Installation: A Simple Crawl (Skip this step if you already run <code>crawl4ai-doctor</code>)</h2>\n<p>Below is a minimal Python script demonstrating a <strong>basic</strong> crawl. It uses our new <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong> for clarity, though no custom settings are passed in this example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.example.com\"</span>,\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])  <span class=\"hljs-comment\"># Show the first 300 characters of extracted text</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Expected</strong> outcome:\n- A headless browser session loads <code>example.com</code>\n- Crawl4AI returns ~300 characters of markdown.<br>\nIf errors occur, rerun <code>crawl4ai-doctor</code> or manually ensure Playwright is installed correctly.</p>\n<hr>\n<h2 id=\"4-advanced-installation-optional\">4. Advanced Installation (Optional)</h2>\n<p><strong>Warning</strong>: Only install these <strong>if you truly need them</strong>. They bring in larger dependencies, including big models, which can increase disk usage and memory load significantly.</p>\n<h3 id=\"41-torch-transformers-or-all\">4.1 Torch, Transformers, or All</h3>\n<ul>\n<li>\n<p><strong>Text Clustering (Torch)</strong><br>\n  </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[torch]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n  Installs PyTorch-based features (e.g., cosine similarity or advanced semantic chunking).<p></p>\n</li>\n<li>\n<p><strong>Transformers</strong><br>\n  </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[transformer]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n  Adds Hugging Face-based summarization or generation strategies.<p></p>\n</li>\n<li>\n<p><strong>All Features</strong><br>\n  </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-css\">pip install crawl4ai<span class=\"hljs-selector-attr\">[all]</span>\ncrawl4ai-setup\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ul>\n<h4 id=\"optional-pre-fetching-models\">(Optional) Pre-Fetching Models</h4>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-undefined\">crawl4ai-download-models\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThis step caches large models locally (if needed). <strong>Only do this</strong> if your workflow requires them.<p></p>\n<hr>\n<h2 id=\"5-docker-experimental\">5. Docker (Experimental)</h2>\n<p>We provide a <strong>temporary</strong> Docker approach for testing. <strong>It’s not stable and may break</strong> with future releases. We plan a major Docker revamp in a future stable version, 2025 Q1. If you still want to try:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">docker pull unclecode/crawl4ai:basic\ndocker run -p 11235:11235 unclecode/crawl4ai:basic\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>You can then make POST requests to <code>http://localhost:11235/crawl</code> to perform crawls. <strong>Production usage</strong> is discouraged until our new Docker approach is ready (planned in Jan or Feb 2025).</p>\n<hr>\n<h2 id=\"6-local-server-mode-legacy\">6. Local Server Mode (Legacy)</h2>\n<p>Some older docs mention running Crawl4AI as a local server. This approach has been <strong>partially replaced</strong> by the new Docker-based prototype and upcoming stable server release. You can experiment, but expect major changes. Official local server instructions will arrive once the new Docker architecture is finalized.</p>\n<hr>\n<h2 id=\"summary\">Summary</h2>\n<p>1. <strong>Install</strong> with <code>pip install crawl4ai</code> and run <code>crawl4ai-setup</code>.\n2. <strong>Diagnose</strong> with <code>crawl4ai-doctor</code> if you see errors.\n3. <strong>Verify</strong> by crawling <code>example.com</code> with minimal <code>BrowserConfig</code> + <code>CrawlerRunConfig</code>.\n4. <strong>Advanced</strong> features (Torch, Transformers) are <strong>optional</strong>—avoid them if you don’t need them (they significantly increase resource usage).\n5. <strong>Docker</strong> is <strong>experimental</strong>—use at your own risk until the stable version is released.\n6. <strong>Local server</strong> references in older docs are largely deprecated; a new solution is in progress.</p>\n<p><strong>Got questions?</strong> Check <a href=\"https://github.com/unclecode/crawl4ai/issues\">GitHub issues</a> for updates or ask the community!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/link-media/",
    "title": "Link & Media - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"link-media\">Link &amp; Media</h1>\n<p>In this tutorial, you’ll learn how to:</p>\n<ol>\n<li>Extract links (internal, external) from crawled pages  </li>\n<li>Filter or exclude specific domains (e.g., social media or custom domains)  </li>\n<li>Access and ma### 3.2 Excluding Images</li>\n</ol>\n<h4 id=\"excluding-external-images\">Excluding External Images</h4>\n<p>If you're dealing with heavy pages or want to skip third-party images (advertisements, for example), you can turn on:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This setting attempts to discard images from outside the primary domain, keeping only those from the site you're crawling.</p>\n<h4 id=\"excluding-all-images\">Excluding All Images</h4>\n<p>If you want to completely remove all images from the page to maximize performance and reduce memory usage, use:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_all_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This setting removes all images very early in the processing pipeline, which significantly improves memory efficiency and processing speed. This is particularly useful when:\n- You don't need image data in your results\n- You're crawling image-heavy pages that cause memory issues\n- You want to focus only on text content\n- You need to maximize crawling speeddata (especially images) in the crawl result<br>\n4. Configure your crawler to exclude or prioritize certain images</p>\n<blockquote>\n<p><strong>Prerequisites</strong><br>\n- You have completed or are familiar with the <a href=\"../simple-crawling/\">AsyncWebCrawler Basics</a> tutorial.<br>\n- You can run Crawl4AI in your environment (Playwright, Python, etc.).</p>\n</blockquote>\n<hr>\n<p>Below is a revised version of the <strong>Link Extraction</strong> and <strong>Media Extraction</strong> sections that includes example data structures showing how links and media items are stored in <code>CrawlResult</code>. Feel free to adjust any field names or descriptions to match your actual output.</p>\n<hr>\n<h2 id=\"1-link-extraction\">1. Link Extraction</h2>\n<h3 id=\"11-resultlinks\">1.1 <code>result.links</code></h3>\n<p>When you call <code>arun()</code> or <code>arun_many()</code> on a URL, Crawl4AI automatically extracts links and stores them in the <code>links</code> field of <code>CrawlResult</code>. By default, the crawler tries to distinguish <strong>internal</strong> links (same domain) from <strong>external</strong> links (different domains).</p>\n<p><strong>Basic Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://www.example.com\"</span>)\n    <span class=\"hljs-keyword\">if</span> result.success:\n        internal_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n        external_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> internal links.\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> external links.\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.media)}</span> media items.\"</span>)\n\n        <span class=\"hljs-comment\"># Each link is typically a dictionary with fields like:</span>\n        <span class=\"hljs-comment\"># { \"href\": \"...\", \"text\": \"...\", \"title\": \"...\", \"base_domain\": \"...\" }</span>\n        <span class=\"hljs-keyword\">if</span> internal_links:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Sample Internal Link:\"</span>, internal_links[<span class=\"hljs-number\">0</span>])\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Structure Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">result.links = {\n  <span class=\"hljs-string\">\"internal\"</span>: [\n    {\n      <span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"https://kidocode.com/\"</span>,\n      <span class=\"hljs-string\">\"text\"</span>: <span class=\"hljs-string\">\"\"</span>,\n      <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"\"</span>,\n      <span class=\"hljs-string\">\"base_domain\"</span>: <span class=\"hljs-string\">\"kidocode.com\"</span>\n    },\n    {\n      <span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"https://kidocode.com/degrees/technology\"</span>,\n      <span class=\"hljs-string\">\"text\"</span>: <span class=\"hljs-string\">\"Technology Degree\"</span>,\n      <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"KidoCode Tech Program\"</span>,\n      <span class=\"hljs-string\">\"base_domain\"</span>: <span class=\"hljs-string\">\"kidocode.com\"</span>\n    },\n    <span class=\"hljs-comment\"># ...</span>\n  ],\n  <span class=\"hljs-string\">\"external\"</span>: [\n    <span class=\"hljs-comment\"># possibly other links leading to third-party sites</span>\n  ]\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ul>\n<li><strong><code>href</code></strong>: The raw hyperlink URL.  </li>\n<li><strong><code>text</code></strong>: The link text (if any) within the <code>&lt;a&gt;</code> tag.  </li>\n<li><strong><code>title</code></strong>: The <code>title</code> attribute of the link (if present).  </li>\n<li><strong><code>base_domain</code></strong>: The domain extracted from <code>href</code>. Helpful for filtering or grouping by domain.</li>\n</ul>\n<hr>\n<h2 id=\"2-advanced-link-head-extraction-scoring\">2. Advanced Link Head Extraction &amp; Scoring</h2>\n<p>Ever wanted to not just extract links, but also get the actual content (title, description, metadata) from those linked pages? And score them for relevance? This is exactly what Link Head Extraction does - it fetches the <code>&lt;head&gt;</code> section from each discovered link and scores them using multiple algorithms.</p>\n<h3 id=\"21-why-link-head-extraction\">2.1 Why Link Head Extraction?</h3>\n<p>When you crawl a page, you get hundreds of links. But which ones are actually valuable? Link Head Extraction solves this by:</p>\n<ol>\n<li><strong>Fetching head content</strong> from each link (title, description, meta tags)</li>\n<li><strong>Scoring links intrinsically</strong> based on URL quality, text relevance, and context</li>\n<li><strong>Scoring links contextually</strong> using BM25 algorithm when you provide a search query</li>\n<li><strong>Combining scores intelligently</strong> to give you a final relevance ranking</li>\n</ol>\n<h3 id=\"22-complete-working-example\">2.2 Complete Working Example</h3>\n<p>Here's a full example you can copy, paste, and run immediately:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LinkPreviewConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_link_heads_example</span>():\n    <span class=\"hljs-string\">\"\"\"\n    Complete example showing link head extraction with scoring.\n    This will crawl a documentation site and extract head content from internal links.\n    \"\"\"</span>\n\n    <span class=\"hljs-comment\"># Configure link head extraction</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Enable link head extraction with detailed configuration</span>\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Extract from internal links</span>\n            include_external=<span class=\"hljs-literal\">False</span>,          <span class=\"hljs-comment\"># Skip external links for this example</span>\n            max_links=<span class=\"hljs-number\">10</span>,                   <span class=\"hljs-comment\"># Limit to 10 links for demo</span>\n            concurrency=<span class=\"hljs-number\">5</span>,                  <span class=\"hljs-comment\"># Process 5 links simultaneously</span>\n            timeout=<span class=\"hljs-number\">10</span>,                     <span class=\"hljs-comment\"># 10 second timeout per link</span>\n            query=<span class=\"hljs-string\">\"API documentation guide\"</span>, <span class=\"hljs-comment\"># Query for contextual scoring</span>\n            score_threshold=<span class=\"hljs-number\">0.3</span>,            <span class=\"hljs-comment\"># Only include links scoring above 0.3</span>\n            verbose=<span class=\"hljs-literal\">True</span>                    <span class=\"hljs-comment\"># Show detailed progress</span>\n        ),\n        <span class=\"hljs-comment\"># Enable intrinsic scoring (URL quality, text relevance)</span>\n        score_links=<span class=\"hljs-literal\">True</span>,\n        <span class=\"hljs-comment\"># Keep output clean</span>\n        only_text=<span class=\"hljs-literal\">True</span>,\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Crawl a documentation site (great for testing)</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✅ Successfully crawled: <span class=\"hljs-subst\">{result.url}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📄 Page title: <span class=\"hljs-subst\">{result.metadata.get(<span class=\"hljs-string\">'title'</span>, <span class=\"hljs-string\">'No title'</span>)}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Access links (now enhanced with head data and scores)</span>\n            internal_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n            external_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n🔗 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(internal_links)}</span> internal links\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🌍 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(external_links)}</span> external links\"</span>)\n\n            <span class=\"hljs-comment\"># Count links with head data</span>\n            links_with_head = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> internal_links \n                             <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"head_data\"</span>) <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>]\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🧠 Links with head data extracted: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(links_with_head)}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Show the top 3 scoring links</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n🏆 Top 3 Links with Full Scoring:\"</span>)\n            <span class=\"hljs-keyword\">for</span> i, link <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(links_with_head[:<span class=\"hljs-number\">3</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{i+<span class=\"hljs-number\">1</span>}</span>. <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Link Text: '<span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'text'</span>, <span class=\"hljs-string\">'No text'</span>)[:<span class=\"hljs-number\">50</span>]}</span>...'\"</span>)\n\n                <span class=\"hljs-comment\"># Show all three score types</span>\n                intrinsic = link.get(<span class=\"hljs-string\">'intrinsic_score'</span>)\n                contextual = link.get(<span class=\"hljs-string\">'contextual_score'</span>) \n                total = link.get(<span class=\"hljs-string\">'total_score'</span>)\n\n                <span class=\"hljs-keyword\">if</span> intrinsic <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📊 Intrinsic Score: <span class=\"hljs-subst\">{intrinsic:<span class=\"hljs-number\">.2</span>f}</span>/10.0 (URL quality &amp; context)\"</span>)\n                <span class=\"hljs-keyword\">if</span> contextual <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   🎯 Contextual Score: <span class=\"hljs-subst\">{contextual:<span class=\"hljs-number\">.3</span>f}</span> (BM25 relevance to query)\"</span>)\n                <span class=\"hljs-keyword\">if</span> total <span class=\"hljs-keyword\">is</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-literal\">None</span>:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   ⭐ Total Score: <span class=\"hljs-subst\">{total:<span class=\"hljs-number\">.3</span>f}</span> (combined final score)\"</span>)\n\n                <span class=\"hljs-comment\"># Show extracted head data</span>\n                head_data = link.get(<span class=\"hljs-string\">\"head_data\"</span>, {})\n                <span class=\"hljs-keyword\">if</span> head_data:\n                    title = head_data.get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"No title\"</span>)\n                    description = head_data.get(<span class=\"hljs-string\">\"meta\"</span>, {}).get(<span class=\"hljs-string\">\"description\"</span>, <span class=\"hljs-string\">\"No description\"</span>)\n\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📰 Title: <span class=\"hljs-subst\">{title[:<span class=\"hljs-number\">60</span>]}</span>...\"</span>)\n                    <span class=\"hljs-keyword\">if</span> description:\n                        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   📝 Description: <span class=\"hljs-subst\">{description[:<span class=\"hljs-number\">80</span>]}</span>...\"</span>)\n\n                    <span class=\"hljs-comment\"># Show extraction status</span>\n                    status = link.get(<span class=\"hljs-string\">\"head_extraction_status\"</span>, <span class=\"hljs-string\">\"unknown\"</span>)\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   ✅ Extraction Status: <span class=\"hljs-subst\">{status}</span>\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"❌ Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-comment\"># Run the example</span>\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(extract_link_heads_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Expected Output:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-yaml\"><span class=\"hljs-string\">✅</span> <span class=\"hljs-attr\">Successfully crawled:</span> <span class=\"hljs-string\">https://docs.python.org/3/</span>\n<span class=\"hljs-string\">📄</span> <span class=\"hljs-attr\">Page title:</span> <span class=\"hljs-number\">3.13</span><span class=\"hljs-number\">.5</span> <span class=\"hljs-string\">Documentation</span>\n<span class=\"hljs-string\">🔗</span> <span class=\"hljs-string\">Found</span> <span class=\"hljs-number\">53</span> <span class=\"hljs-string\">internal</span> <span class=\"hljs-string\">links</span>\n<span class=\"hljs-string\">🌍</span> <span class=\"hljs-string\">Found</span> <span class=\"hljs-number\">1</span> <span class=\"hljs-string\">external</span> <span class=\"hljs-string\">links</span>\n<span class=\"hljs-string\">🧠</span> <span class=\"hljs-attr\">Links with head data extracted:</span> <span class=\"hljs-number\">10</span>\n\n<span class=\"hljs-string\">🏆</span> <span class=\"hljs-attr\">Top 3 Links with Full Scoring:</span>\n\n<span class=\"hljs-number\">1</span><span class=\"hljs-string\">.</span> <span class=\"hljs-string\">https://docs.python.org/3.15/</span>\n   <span class=\"hljs-attr\">Link Text:</span> <span class=\"hljs-string\">'Python 3.15 (in development)...'</span>\n   <span class=\"hljs-string\">📊</span> <span class=\"hljs-attr\">Intrinsic Score:</span> <span class=\"hljs-number\">4.17</span><span class=\"hljs-string\">/10.0</span> <span class=\"hljs-string\">(URL</span> <span class=\"hljs-string\">quality</span> <span class=\"hljs-string\">&amp;</span> <span class=\"hljs-string\">context)</span>\n   <span class=\"hljs-string\">🎯</span> <span class=\"hljs-attr\">Contextual Score:</span> <span class=\"hljs-number\">1.000</span> <span class=\"hljs-string\">(BM25</span> <span class=\"hljs-string\">relevance</span> <span class=\"hljs-string\">to</span> <span class=\"hljs-string\">query)</span>\n   <span class=\"hljs-string\">⭐</span> <span class=\"hljs-attr\">Total Score:</span> <span class=\"hljs-number\">5.917</span> <span class=\"hljs-string\">(combined</span> <span class=\"hljs-string\">final</span> <span class=\"hljs-string\">score)</span>\n   <span class=\"hljs-string\">📰</span> <span class=\"hljs-attr\">Title:</span> <span class=\"hljs-number\">3.15</span><span class=\"hljs-string\">.0a0</span> <span class=\"hljs-string\">Documentation...</span>\n   <span class=\"hljs-string\">📝</span> <span class=\"hljs-attr\">Description:</span> <span class=\"hljs-string\">The</span> <span class=\"hljs-string\">official</span> <span class=\"hljs-string\">Python</span> <span class=\"hljs-string\">documentation...</span>\n   <span class=\"hljs-string\">✅</span> <span class=\"hljs-attr\">Extraction Status:</span> <span class=\"hljs-string\">valid</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"23-configuration-deep-dive\">2.3 Configuration Deep Dive</h3>\n<p>The <code>LinkPreviewConfig</code> class supports these options:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LinkPreviewConfig\n\nlink_preview_config = LinkPreviewConfig(\n    <span class=\"hljs-comment\"># BASIC SETTINGS</span>\n    verbose=<span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Show detailed logs (recommended for learning)</span>\n\n    <span class=\"hljs-comment\"># LINK FILTERING</span>\n    include_internal=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Include same-domain links</span>\n    include_external=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Include different-domain links</span>\n    max_links=<span class=\"hljs-number\">50</span>,                   <span class=\"hljs-comment\"># Maximum links to process (prevents overload)</span>\n\n    <span class=\"hljs-comment\"># PATTERN FILTERING</span>\n    include_patterns=[               <span class=\"hljs-comment\"># Only process links matching these patterns</span>\n        <span class=\"hljs-string\">\"*/docs/*\"</span>, \n        <span class=\"hljs-string\">\"*/api/*\"</span>, \n        <span class=\"hljs-string\">\"*/reference/*\"</span>\n    ],\n    exclude_patterns=[               <span class=\"hljs-comment\"># Skip links matching these patterns</span>\n        <span class=\"hljs-string\">\"*/login*\"</span>,\n        <span class=\"hljs-string\">\"*/admin*\"</span>\n    ],\n\n    <span class=\"hljs-comment\"># PERFORMANCE SETTINGS</span>\n    concurrency=<span class=\"hljs-number\">10</span>,                  <span class=\"hljs-comment\"># How many links to process simultaneously</span>\n    timeout=<span class=\"hljs-number\">5</span>,                      <span class=\"hljs-comment\"># Seconds to wait per link</span>\n\n    <span class=\"hljs-comment\"># RELEVANCE SCORING</span>\n    query=<span class=\"hljs-string\">\"machine learning API\"</span>,    <span class=\"hljs-comment\"># Query for BM25 contextual scoring</span>\n    score_threshold=<span class=\"hljs-number\">0.3</span>,            <span class=\"hljs-comment\"># Only include links above this score</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"24-understanding-the-three-score-types\">2.4 Understanding the Three Score Types</h3>\n<p>Each extracted link gets three different scores:</p>\n<h4 id=\"1-intrinsic-score-0-10-url-and-content-quality\">1. <strong>Intrinsic Score (0-10)</strong> - URL and Content Quality</h4>\n<p>Based on URL structure, link text quality, and page context:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># High intrinsic score indicators:</span>\n<span class=\"hljs-comment\"># ✅ Clean URL structure (docs.python.org/api/reference)</span>\n<span class=\"hljs-comment\"># ✅ Meaningful link text (\"API Reference Guide\")</span>\n<span class=\"hljs-comment\"># ✅ Relevant to page context</span>\n<span class=\"hljs-comment\"># ✅ Not buried deep in navigation</span>\n\n<span class=\"hljs-comment\"># Low intrinsic score indicators:</span>\n<span class=\"hljs-comment\"># ❌ Random URLs (site.com/x7f9g2h)</span>\n<span class=\"hljs-comment\"># ❌ No link text or generic text (\"Click here\")</span>\n<span class=\"hljs-comment\"># ❌ Unrelated to page content</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"2-contextual-score-0-1-bm25-relevance-to-query\">2. <strong>Contextual Score (0-1)</strong> - BM25 Relevance to Query</h4>\n<p>Only available when you provide a <code>query</code>. Uses BM25 algorithm against head content:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-shell\"><span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">Example: query = <span class=\"hljs-string\">\"machine learning tutorial\"</span></span>\n<span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">High contextual score: Link to <span class=\"hljs-string\">\"Complete Machine Learning Guide\"</span></span>\n<span class=\"hljs-meta prompt_\"># </span><span class=\"language-bash\">Low contextual score: Link to <span class=\"hljs-string\">\"Privacy Policy\"</span></span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"3-total-score-smart-combination\">3. <strong>Total Score</strong> - Smart Combination</h4>\n<p>Intelligently combines intrinsic and contextual scores with fallbacks:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># When both scores available: (intrinsic * 0.3) + (contextual * 0.7)</span>\n<span class=\"hljs-comment\"># When only intrinsic: uses intrinsic score</span>\n<span class=\"hljs-comment\"># When only contextual: uses contextual score</span>\n<span class=\"hljs-comment\"># When neither: not calculated</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"25-practical-use-cases\">2.5 Practical Use Cases</h3>\n<h4 id=\"use-case-1-research-assistant\">Use Case 1: Research Assistant</h4>\n<p>Find the most relevant documentation pages:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">research_assistant</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            include_external=<span class=\"hljs-literal\">True</span>,\n            include_patterns=[<span class=\"hljs-string\">\"*/docs/*\"</span>, <span class=\"hljs-string\">\"*/tutorial/*\"</span>, <span class=\"hljs-string\">\"*/guide/*\"</span>],\n            query=<span class=\"hljs-string\">\"machine learning neural networks\"</span>,\n            max_links=<span class=\"hljs-number\">20</span>,\n            score_threshold=<span class=\"hljs-number\">0.5</span>,  <span class=\"hljs-comment\"># Only high-relevance links</span>\n            verbose=<span class=\"hljs-literal\">True</span>\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://scikit-learn.org/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># Get high-scoring links</span>\n            good_links = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n                         <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"total_score\"</span>, <span class=\"hljs-number\">0</span>) &gt; <span class=\"hljs-number\">0.7</span>]\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🎯 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(good_links)}</span> highly relevant links:\"</span>)\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> good_links[:<span class=\"hljs-number\">5</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"⭐ <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'total_score'</span>]:<span class=\"hljs-number\">.3</span>f}</span> - <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   <span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'head_data'</span>, {}</span>).get('title', 'No title')}\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"use-case-2-content-discovery\">Use Case 2: Content Discovery</h4>\n<p>Find all API endpoints and references:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">api_discovery</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            include_patterns=[<span class=\"hljs-string\">\"*/api/*\"</span>, <span class=\"hljs-string\">\"*/reference/*\"</span>],\n            exclude_patterns=[<span class=\"hljs-string\">\"*/deprecated/*\"</span>],\n            max_links=<span class=\"hljs-number\">100</span>,\n            concurrency=<span class=\"hljs-number\">15</span>,\n            verbose=<span class=\"hljs-literal\">False</span>  <span class=\"hljs-comment\"># Clean output</span>\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://docs.example-api.com/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            api_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n\n            <span class=\"hljs-comment\"># Group by endpoint type</span>\n            endpoints = {}\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> api_links:\n                <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">\"head_data\"</span>):\n                    title = link[<span class=\"hljs-string\">\"head_data\"</span>].get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"\"</span>)\n                    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">\"GET\"</span> <span class=\"hljs-keyword\">in</span> title:\n                        endpoints.setdefault(<span class=\"hljs-string\">\"GET\"</span>, []).append(link)\n                    <span class=\"hljs-keyword\">elif</span> <span class=\"hljs-string\">\"POST\"</span> <span class=\"hljs-keyword\">in</span> title:\n                        endpoints.setdefault(<span class=\"hljs-string\">\"POST\"</span>, []).append(link)\n\n            <span class=\"hljs-keyword\">for</span> method, links <span class=\"hljs-keyword\">in</span> endpoints.items():\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{method}</span> Endpoints (<span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(links)}</span>):\"</span>)\n                <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links[:<span class=\"hljs-number\">3</span>]:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  • <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"use-case-3-link-quality-analysis\">Use Case 3: Link Quality Analysis</h4>\n<p>Analyze website structure and content quality:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">quality_analysis</span>():\n    config = CrawlerRunConfig(\n        link_preview_config=LinkPreviewConfig(\n            include_internal=<span class=\"hljs-literal\">True</span>,\n            max_links=<span class=\"hljs-number\">200</span>,\n            concurrency=<span class=\"hljs-number\">20</span>,\n        ),\n        score_links=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://your-website.com/\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n\n            <span class=\"hljs-comment\"># Analyze intrinsic scores</span>\n            scores = [link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>) <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links]\n            avg_score = <span class=\"hljs-built_in\">sum</span>(scores) / <span class=\"hljs-built_in\">len</span>(scores) <span class=\"hljs-keyword\">if</span> scores <span class=\"hljs-keyword\">else</span> <span class=\"hljs-number\">0</span>\n\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📊 Link Quality Analysis:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Average intrinsic score: <span class=\"hljs-subst\">{avg_score:<span class=\"hljs-number\">.2</span>f}</span>/10.0\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   High quality links (&gt;7.0): <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>([s <span class=\"hljs-keyword\">for</span> s <span class=\"hljs-keyword\">in</span> scores <span class=\"hljs-keyword\">if</span> s &gt; <span class=\"hljs-number\">7.0</span>])}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Low quality links (&lt;3.0): <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>([s <span class=\"hljs-keyword\">for</span> s <span class=\"hljs-keyword\">in</span> scores <span class=\"hljs-keyword\">if</span> s &lt; <span class=\"hljs-number\">3.0</span>])}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Find problematic links</span>\n            bad_links = [link <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> links \n                        <span class=\"hljs-keyword\">if</span> link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>) &lt; <span class=\"hljs-number\">2.0</span>]\n\n            <span class=\"hljs-keyword\">if</span> bad_links:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n⚠️  Links needing attention:\"</span>)\n                <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> bad_links[:<span class=\"hljs-number\">5</span>]:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span> (score: <span class=\"hljs-subst\">{link.get(<span class=\"hljs-string\">'intrinsic_score'</span>, <span class=\"hljs-number\">0</span>):<span class=\"hljs-number\">.1</span>f}</span>)\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"26-performance-tips\">2.6 Performance Tips</h3>\n<ol>\n<li><strong>Start Small</strong>: Begin with <code>max_links: 10</code> to understand the feature</li>\n<li><strong>Use Patterns</strong>: Filter with <code>include_patterns</code> to focus on relevant sections</li>\n<li><strong>Adjust Concurrency</strong>: Higher concurrency = faster but more resource usage</li>\n<li><strong>Set Timeouts</strong>: Use <code>timeout: 5</code> to prevent hanging on slow sites</li>\n<li><strong>Use Score Thresholds</strong>: Filter out low-quality links with <code>score_threshold</code></li>\n</ol>\n<h3 id=\"27-troubleshooting\">2.7 Troubleshooting</h3>\n<p><strong>No head data extracted?</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Check your configuration:</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    link_preview_config<span class=\"hljs-punctuation\">=</span>LinkPreviewConfig<span class=\"hljs-punctuation\">(</span>\n        verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>   <span class=\"hljs-comment\"># ← Enable to see what's happening</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Scores showing as None?</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Make sure scoring is enabled:</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    score_links<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># ← Enable intrinsic scoring</span>\n    link_preview_config<span class=\"hljs-punctuation\">=</span>LinkPreviewConfig<span class=\"hljs-punctuation\">(</span>\n        <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"your search terms\"</span>  <span class=\"hljs-comment\"># ← For contextual scoring</span>\n    <span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Process taking too long?</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Optimize performance:</span>\nlink_preview_config = LinkPreviewConfig(\n    max_links=20,      <span class=\"hljs-comment\"># ← Reduce number</span>\n    concurrency=10,    <span class=\"hljs-comment\"># ← Increase parallelism</span>\n    <span class=\"hljs-built_in\">timeout</span>=3,         <span class=\"hljs-comment\"># ← Shorter timeout</span>\n    include_patterns=[<span class=\"hljs-string\">\"*/important/*\"</span>]  <span class=\"hljs-comment\"># ← Focus on key areas</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<hr>\n<h2 id=\"3-domain-filtering\">3. Domain Filtering</h2>\n<p>Some websites contain hundreds of third-party or affiliate links. You can filter out certain domains at <strong>crawl time</strong> by configuring the crawler. The most relevant parameters in <code>CrawlerRunConfig</code> are:</p>\n<ul>\n<li><strong><code>exclude_external_links</code></strong>: If <code>True</code>, discard any link pointing outside the root domain.  </li>\n<li><strong><code>exclude_social_media_domains</code></strong>: Provide a list of social media platforms (e.g., <code>[\"facebook.com\", \"twitter.com\"]</code>) to exclude from your crawl.  </li>\n<li><strong><code>exclude_social_media_links</code></strong>: If <code>True</code>, automatically skip known social platforms.  </li>\n<li><strong><code>exclude_domains</code></strong>: Provide a list of custom domains you want to exclude (e.g., <code>[\"spammyads.com\", \"tracker.net\"]</code>).</li>\n</ul>\n<h3 id=\"31-example-excluding-external-social-media-links\">3.1 Example: Excluding External &amp; Social Media Links</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    crawler_cfg = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,          <span class=\"hljs-comment\"># No links outside primary domain</span>\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>       <span class=\"hljs-comment\"># Skip recognized social media domains</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            <span class=\"hljs-string\">\"https://www.example.com\"</span>,\n            config=crawler_cfg\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Crawled:\"</span>, result.url)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal links count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"External links count:\"</span>, <span class=\"hljs-built_in\">len</span>(result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])))  \n            <span class=\"hljs-comment\"># Likely zero external links in this scenario</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR]\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"32-example-excluding-specific-domains\">3.2 Example: Excluding Specific Domains</h3>\n<p>If you want to let external links in, but specifically exclude a domain (e.g., <code>suspiciousads.com</code>), do this:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">crawler_cfg = CrawlerRunConfig(\n    exclude_domains=[<span class=\"hljs-string\">\"suspiciousads.com\"</span>]\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This approach is handy when you still want external links but need to block certain sites you consider spammy.</p>\n<hr>\n<h2 id=\"4-media-extraction\">4. Media Extraction</h2>\n<h3 id=\"41-accessing-resultmedia\">4.1 Accessing <code>result.media</code></h3>\n<p>By default, Crawl4AI collects images, audio and video URLs it finds on the page. These are stored in <code>result.media</code>, a dictionary keyed by media type (e.g., <code>images</code>, <code>videos</code>, <code>audio</code>).\n<strong>Note: Tables have been moved from <code>result.media[\"tables\"]</code> to the new <code>result.tables</code> format for better organization and direct access.</strong></p>\n<p><strong>Basic Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">if</span> result.success:\n    <span class=\"hljs-comment\"># Get images</span>\n    images_info = result.media.get(<span class=\"hljs-string\">\"images\"</span>, [])\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(images_info)}</span> images in total.\"</span>)\n    <span class=\"hljs-keyword\">for</span> i, img <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(images_info[:<span class=\"hljs-number\">3</span>]):  <span class=\"hljs-comment\"># Inspect just the first 3</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[Image <span class=\"hljs-subst\">{i}</span>] URL: <span class=\"hljs-subst\">{img[<span class=\"hljs-string\">'src'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Alt text: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'alt'</span>, <span class=\"hljs-string\">''</span>)}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Score: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'score'</span>)}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"           Description: <span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'desc'</span>, <span class=\"hljs-string\">''</span>)}</span>\\n\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Structure Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">result.media = {\n  <span class=\"hljs-string\">\"images\"</span>: [\n    {\n      <span class=\"hljs-string\">\"src\"</span>: <span class=\"hljs-string\">\"https://cdn.prod.website-files.com/.../Group%2089.svg\"</span>,\n      <span class=\"hljs-string\">\"alt\"</span>: <span class=\"hljs-string\">\"coding school for kids\"</span>,\n      <span class=\"hljs-string\">\"desc\"</span>: <span class=\"hljs-string\">\"Trial Class Degrees degrees All Degrees AI Degree Technology ...\"</span>,\n      <span class=\"hljs-string\">\"score\"</span>: <span class=\"hljs-number\">3</span>,\n      <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"image\"</span>,\n      <span class=\"hljs-string\">\"group_id\"</span>: <span class=\"hljs-number\">0</span>,\n      <span class=\"hljs-string\">\"format\"</span>: <span class=\"hljs-literal\">None</span>,\n      <span class=\"hljs-string\">\"width\"</span>: <span class=\"hljs-literal\">None</span>,\n      <span class=\"hljs-string\">\"height\"</span>: <span class=\"hljs-literal\">None</span>\n    },\n    <span class=\"hljs-comment\"># ...</span>\n  ],\n  <span class=\"hljs-string\">\"videos\"</span>: [\n    <span class=\"hljs-comment\"># Similar structure but with video-specific fields</span>\n  ],\n  <span class=\"hljs-string\">\"audio\"</span>: [\n    <span class=\"hljs-comment\"># Similar structure but with audio-specific fields</span>\n  ],\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Depending on your Crawl4AI version or scraping strategy, these dictionaries can include fields like:</p>\n<ul>\n<li><strong><code>src</code></strong>: The media URL (e.g., image source)  </li>\n<li><strong><code>alt</code></strong>: The alt text for images (if present)  </li>\n<li><strong><code>desc</code></strong>: A snippet of nearby text or a short description (optional)  </li>\n<li><strong><code>score</code></strong>: A heuristic relevance score if you’re using content-scoring features  </li>\n<li><strong><code>width</code></strong>, <strong><code>height</code></strong>: If the crawler detects dimensions for the image/video  </li>\n<li><strong><code>type</code></strong>: Usually <code>\"image\"</code>, <code>\"video\"</code>, or <code>\"audio\"</code>  </li>\n<li><strong><code>group_id</code></strong>: If you’re grouping related media items, the crawler might assign an ID  </li>\n</ul>\n<p>With these details, you can easily filter out or focus on certain images (for instance, ignoring images with very low scores or a different domain), or gather metadata for analytics.</p>\n<h3 id=\"42-excluding-external-images\">4.2 Excluding External Images</h3>\n<p>If you’re dealing with heavy pages or want to skip third-party images (advertisements, for example), you can turn on:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">crawler_cfg <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    exclude_external_images<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This setting attempts to discard images from outside the primary domain, keeping only those from the site you’re crawling.</p>\n<h3 id=\"43-additional-media-config\">4.3 Additional Media Config</h3>\n<ul>\n<li><strong><code>screenshot</code></strong>: Set to <code>True</code> if you want a full-page screenshot stored as <code>base64</code> in <code>result.screenshot</code>.  </li>\n<li><strong><code>pdf</code></strong>: Set to <code>True</code> if you want a PDF version of the page in <code>result.pdf</code>.  </li>\n<li><strong><code>capture_mhtml</code></strong>: Set to <code>True</code> if you want an MHTML snapshot of the page in <code>result.mhtml</code>. This format preserves the entire web page with all its resources (CSS, images, scripts) in a single file, making it perfect for archiving or offline viewing.</li>\n<li><strong><code>wait_for_images</code></strong>: If <code>True</code>, attempts to wait until images are fully loaded before final extraction.</li>\n</ul>\n<h4 id=\"example-capturing-page-as-mhtml\">Example: Capturing Page as MHTML</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    crawler_cfg = CrawlerRunConfig(\n        capture_mhtml=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable MHTML capture</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=crawler_cfg)\n\n        <span class=\"hljs-keyword\">if</span> result.success <span class=\"hljs-keyword\">and</span> result.mhtml:\n            <span class=\"hljs-comment\"># Save the MHTML snapshot to a file</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"example.mhtml\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                f.write(result.mhtml)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"MHTML snapshot saved to example.mhtml\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Failed to capture MHTML:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The MHTML format is particularly useful because:\n- It captures the complete page state including all resources\n- It can be opened in most modern browsers for offline viewing\n- It preserves the page exactly as it appeared during crawling\n- It's a single file, making it easy to store and transfer</p>\n<hr>\n<h2 id=\"5-putting-it-all-together-link-media-filtering\">5. Putting It All Together: Link &amp; Media Filtering</h2>\n<p>Here’s a combined example demonstrating how to filter out external links, skip certain domains, and exclude external images:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Suppose we want to keep only internal links, remove certain domains, </span>\n    <span class=\"hljs-comment\"># and discard external images from the final crawl data.</span>\n    crawler_cfg = CrawlerRunConfig(\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n        exclude_domains=[<span class=\"hljs-string\">\"spammyads.com\"</span>],\n        exclude_social_media_links=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># skip Twitter, Facebook, etc.</span>\n        exclude_external_images=<span class=\"hljs-literal\">True</span>,      <span class=\"hljs-comment\"># keep only images from main domain</span>\n        wait_for_images=<span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># ensure images are loaded</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://www.example.com\"</span>, config=crawler_cfg)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[OK] Crawled:\"</span>, result.url)\n\n            <span class=\"hljs-comment\"># 1. Links</span>\n            in_links = result.links.get(<span class=\"hljs-string\">\"internal\"</span>, [])\n            ext_links = result.links.get(<span class=\"hljs-string\">\"external\"</span>, [])\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Internal link count:\"</span>, <span class=\"hljs-built_in\">len</span>(in_links))\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"External link count:\"</span>, <span class=\"hljs-built_in\">len</span>(ext_links))  <span class=\"hljs-comment\"># should be zero with exclude_external_links=True</span>\n\n            <span class=\"hljs-comment\"># 2. Images</span>\n            images = result.media.get(<span class=\"hljs-string\">\"images\"</span>, [])\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Images found:\"</span>, <span class=\"hljs-built_in\">len</span>(images))\n\n            <span class=\"hljs-comment\"># Let's see a snippet of these images</span>\n            <span class=\"hljs-keyword\">for</span> i, img <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(images[:<span class=\"hljs-number\">3</span>]):\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  - <span class=\"hljs-subst\">{img[<span class=\"hljs-string\">'src'</span>]}</span> (alt=<span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'alt'</span>,<span class=\"hljs-string\">''</span>)}</span>, score=<span class=\"hljs-subst\">{img.get(<span class=\"hljs-string\">'score'</span>,<span class=\"hljs-string\">'N/A'</span>)}</span>)\"</span>)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"[ERROR] Failed to crawl. Reason:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"6-common-pitfalls-tips\">6. Common Pitfalls &amp; Tips</h2>\n<p>1. <strong>Conflicting Flags</strong>:<br>\n   - <code>exclude_external_links=True</code> but then also specifying <code>exclude_social_media_links=True</code> is typically fine, but understand that the first setting already discards <em>all</em> external links. The second becomes somewhat redundant.<br>\n   - <code>exclude_external_images=True</code> but want to keep some external images? Currently no partial domain-based setting for images, so you might need a custom approach or hook logic.</p>\n<p>2. <strong>Relevancy Scores</strong>:<br>\n   - If your version of Crawl4AI or your scraping strategy includes an <code>img[\"score\"]</code>, it’s typically a heuristic based on size, position, or content analysis. Evaluate carefully if you rely on it.</p>\n<p>3. <strong>Performance</strong>:<br>\n   - Excluding certain domains or external images can speed up your crawl, especially for large, media-heavy pages.<br>\n   - If you want a “full” link map, do <em>not</em> exclude them. Instead, you can post-filter in your own code.</p>\n<p>4. <strong>Social Media Lists</strong>:<br>\n   - <code>exclude_social_media_links=True</code> typically references an internal list of known social domains like Facebook, Twitter, LinkedIn, etc. If you need to add or remove from that list, look for library settings or a local config file (depending on your version).</p>\n<hr>\n<p><strong>That’s it for Link &amp; Media Analysis!</strong> You’re now equipped to filter out unwanted sites and zero in on the images and videos that matter for your project.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/llmtxt/",
    "title": "Llmtxt - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <p>I</p><div class=\"llmtxt-container\"><p></p>\n\n</div>\n\n\n\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/local-files/",
    "title": "Local Files & Raw HTML - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"prefix-based-input-handling-in-crawl4ai\">Prefix-Based Input Handling in Crawl4AI</h1>\n<p>This guide will walk you through using the Crawl4AI library to crawl web pages, local HTML files, and raw HTML strings. We'll demonstrate these capabilities using a Wikipedia page as an example.</p>\n<h2 id=\"crawling-a-web-url\">Crawling a Web URL</h2>\n<p>To crawl a live web page, provide the URL starting with <code>http://</code> or <code>https://</code>, using a <code>CrawlerRunConfig</code> object:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_web</span>():\n    config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://en.wikipedia.org/wiki/apple\"</span>, \n            config=config\n        )\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown Content:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\nasyncio.run(crawl_web())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"crawling-a-local-html-file\">Crawling a Local HTML File</h2>\n<p>To crawl a local HTML file, prefix the file path with <code>file://</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_local_file</span>():\n    local_file_path = <span class=\"hljs-string\">\"/path/to/apple.html\"</span>  <span class=\"hljs-comment\"># Replace with your file path</span>\n    file_url = <span class=\"hljs-string\">f\"file://<span class=\"hljs-subst\">{local_file_path}</span>\"</span>\n    config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=file_url, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown Content from Local File:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl local file: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\nasyncio.run(crawl_local_file())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"crawling-raw-html-content\">Crawling Raw HTML Content</h2>\n<p>To crawl raw HTML content, prefix the HTML string with <code>raw:</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_raw_html</span>():\n    raw_html = <span class=\"hljs-string\">\"&lt;html&gt;&lt;body&gt;&lt;h1&gt;Hello, World!&lt;/h1&gt;&lt;/body&gt;&lt;/html&gt;\"</span>\n    raw_html_url = <span class=\"hljs-string\">f\"raw:<span class=\"hljs-subst\">{raw_html}</span>\"</span>\n    config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=raw_html_url, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown Content from Raw HTML:\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl raw HTML: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\nasyncio.run(crawl_raw_html())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h1 id=\"complete-example\">Complete Example</h1>\n<p>Below is a comprehensive script that:</p>\n<ol>\n<li>Crawls the Wikipedia page for \"Apple.\"</li>\n<li>Saves the HTML content to a local file (<code>apple.html</code>).</li>\n<li>Crawls the local HTML file and verifies the markdown length matches the original crawl.</li>\n<li>Crawls the raw HTML content from the saved file and verifies consistency.</li>\n</ol>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> sys\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> pathlib <span class=\"hljs-keyword\">import</span> Path\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CacheMode, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    wikipedia_url = <span class=\"hljs-string\">\"https://en.wikipedia.org/wiki/apple\"</span>\n    script_dir = Path(__file__).parent\n    html_file_path = script_dir / <span class=\"hljs-string\">\"apple.html\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Step 1: Crawl the Web URL</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n=== Step 1: Crawling the Wikipedia URL ===\"</span>)\n        web_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=wikipedia_url, config=web_config)\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl <span class=\"hljs-subst\">{wikipedia_url}</span>: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(html_file_path, <span class=\"hljs-string\">'w'</span>, encoding=<span class=\"hljs-string\">'utf-8'</span>) <span class=\"hljs-keyword\">as</span> f:\n            f.write(result.html)\n        web_crawl_length = <span class=\"hljs-built_in\">len</span>(result.markdown)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Length of markdown from web crawl: <span class=\"hljs-subst\">{web_crawl_length}</span>\\n\"</span>)\n\n        <span class=\"hljs-comment\"># Step 2: Crawl from the Local HTML File</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"=== Step 2: Crawling from the Local HTML File ===\"</span>)\n        file_url = <span class=\"hljs-string\">f\"file://<span class=\"hljs-subst\">{html_file_path.resolve()}</span>\"</span>\n        file_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n        local_result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=file_url, config=file_config)\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> local_result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl local file <span class=\"hljs-subst\">{file_url}</span>: <span class=\"hljs-subst\">{local_result.error_message}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        local_crawl_length = <span class=\"hljs-built_in\">len</span>(local_result.markdown)\n        <span class=\"hljs-keyword\">assert</span> web_crawl_length == local_crawl_length, <span class=\"hljs-string\">\"Markdown length mismatch\"</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"✅ Markdown length matches between web and local file crawl.\\n\"</span>)\n\n        <span class=\"hljs-comment\"># Step 3: Crawl Using Raw HTML Content</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"=== Step 3: Crawling Using Raw HTML Content ===\"</span>)\n        <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(html_file_path, <span class=\"hljs-string\">'r'</span>, encoding=<span class=\"hljs-string\">'utf-8'</span>) <span class=\"hljs-keyword\">as</span> f:\n            raw_html_content = f.read()\n        raw_html_url = <span class=\"hljs-string\">f\"raw:<span class=\"hljs-subst\">{raw_html_content}</span>\"</span>\n        raw_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)\n        raw_result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=raw_html_url, config=raw_config)\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> raw_result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Failed to crawl raw HTML content: <span class=\"hljs-subst\">{raw_result.error_message}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        raw_crawl_length = <span class=\"hljs-built_in\">len</span>(raw_result.markdown)\n        <span class=\"hljs-keyword\">assert</span> web_crawl_length == raw_crawl_length, <span class=\"hljs-string\">\"Markdown length mismatch\"</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"✅ Markdown length matches between web and raw HTML crawl.\\n\"</span>)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"All tests passed successfully!\"</span>)\n    <span class=\"hljs-keyword\">if</span> html_file_path.exists():\n        os.remove(html_file_path)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h1 id=\"conclusion\">Conclusion</h1>\n<p>With the unified <code>url</code> parameter and prefix-based handling in <strong>Crawl4AI</strong>, you can seamlessly handle web URLs, local HTML files, and raw HTML content. Use <code>CrawlerRunConfig</code> for flexible and consistent configuration in all scenarios.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/markdown-generation/",
    "title": "Markdown Generation - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"markdown-generation-basics\">Markdown Generation Basics</h1>\n<p>One of Crawl4AI’s core features is generating <strong>clean, structured markdown</strong> from web pages. Originally built to solve the problem of extracting only the “actual” content and discarding boilerplate or noise, Crawl4AI’s markdown system remains one of its biggest draws for AI workflows.</p>\n<p>In this tutorial, you’ll learn:</p>\n<ol>\n<li>How to configure the <strong>Default Markdown Generator</strong>  </li>\n<li>How <strong>content filters</strong> (BM25 or Pruning) help you refine markdown and discard junk  </li>\n<li>The difference between raw markdown (<code>result.markdown</code>) and filtered markdown (<code>fit_markdown</code>)  </li>\n</ol>\n<blockquote>\n<p><strong>Prerequisites</strong><br>\n- You’ve completed or read <a href=\"../simple-crawling/\">AsyncWebCrawler Basics</a> to understand how to run a simple crawl.<br>\n- You know how to configure <code>CrawlerRunConfig</code>.</p>\n</blockquote>\n<hr>\n<h2 id=\"1-quick-example\">1. Quick Example</h2>\n<p>Here’s a minimal code snippet that uses the <strong>DefaultMarkdownGenerator</strong> with no additional filtering:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        markdown_generator=DefaultMarkdownGenerator()\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown Output:\\n\"</span>)\n            <span class=\"hljs-built_in\">print</span>(result.markdown)  <span class=\"hljs-comment\"># The unfiltered markdown from the page</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What’s happening?</strong><br>\n- <code>CrawlerRunConfig( markdown_generator = DefaultMarkdownGenerator() )</code> instructs Crawl4AI to convert the final HTML into markdown at the end of each crawl.<br>\n- The resulting markdown is accessible via <code>result.markdown</code>.</p>\n<hr>\n<h2 id=\"2-how-markdown-generation-works\">2. How Markdown Generation Works</h2>\n<h3 id=\"21-html-to-text-conversion-forked-modified\">2.1 HTML-to-Text Conversion (Forked &amp; Modified)</h3>\n<p>Under the hood, <strong>DefaultMarkdownGenerator</strong> uses a specialized HTML-to-text approach that:</p>\n<ul>\n<li>Preserves headings, code blocks, bullet points, etc.  </li>\n<li>Removes extraneous tags (scripts, styles) that don’t add meaningful content.  </li>\n<li>Can optionally generate references for links or skip them altogether.</li>\n</ul>\n<p>A set of <strong>options</strong> (passed as a dict) allows you to customize precisely how HTML converts to markdown. These map to standard html2text-like configuration plus your own enhancements (e.g., ignoring internal links, preserving certain tags verbatim, or adjusting line widths).</p>\n<h3 id=\"22-link-citations-references\">2.2 Link Citations &amp; References</h3>\n<p>By default, the generator can convert <code>&lt;a href=\"...\"&gt;</code> elements into <code>[text][1]</code> citations, then place the actual links at the bottom of the document. This is handy for research workflows that demand references in a structured manner.</p>\n<h3 id=\"23-optional-content-filters\">2.3 Optional Content Filters</h3>\n<p>Before or after the HTML-to-Markdown step, you can apply a <strong>content filter</strong> (like BM25 or Pruning) to reduce noise and produce a “fit_markdown”—a heavily pruned version focusing on the page’s main text. We’ll cover these filters shortly.</p>\n<hr>\n<h2 id=\"3-configuring-the-default-markdown-generator\">3. Configuring the Default Markdown Generator</h2>\n<p>You can tweak the output by passing an <code>options</code> dict to <code>DefaultMarkdownGenerator</code>. For example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Example: ignore all links, don't escape HTML, and wrap text at 80 characters</span>\n    md_generator = DefaultMarkdownGenerator(\n        options={\n            <span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>,\n            <span class=\"hljs-string\">\"escape_html\"</span>: <span class=\"hljs-literal\">False</span>,\n            <span class=\"hljs-string\">\"body_width\"</span>: <span class=\"hljs-number\">80</span>\n        }\n    )\n\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/docs\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown:\\n\"</span>, result.markdown[:<span class=\"hljs-number\">500</span>])  <span class=\"hljs-comment\"># Just a snippet</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    <span class=\"hljs-keyword\">import</span> asyncio\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Some commonly used <code>options</code>:</p>\n<ul>\n<li><strong><code>ignore_links</code></strong> (bool): Whether to remove all hyperlinks in the final markdown.  </li>\n<li><strong><code>ignore_images</code></strong> (bool): Remove all <code>![image]()</code> references.  </li>\n<li><strong><code>escape_html</code></strong> (bool): Turn HTML entities into text (default is often <code>True</code>).  </li>\n<li><strong><code>body_width</code></strong> (int): Wrap text at N characters. <code>0</code> or <code>None</code> means no wrapping.  </li>\n<li><strong><code>skip_internal_links</code></strong> (bool): If <code>True</code>, omit <code>#localAnchors</code> or internal links referencing the same page.  </li>\n<li><strong><code>include_sup_sub</code></strong> (bool): Attempt to handle <code>&lt;sup&gt;</code> / <code>&lt;sub&gt;</code> in a more readable way.</li>\n</ul>\n<h2 id=\"4-selecting-the-html-source-for-markdown-generation\">4. Selecting the HTML Source for Markdown Generation</h2>\n<p>The <code>content_source</code> parameter allows you to control which HTML content is used as input for markdown generation. This gives you flexibility in how the HTML is processed before conversion to markdown.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Option 1: Use the raw HTML directly from the webpage (before any processing)</span>\n    raw_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"raw_html\"</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Option 2: Use the cleaned HTML (after scraping strategy processing - default)</span>\n    cleaned_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"cleaned_html\"</span>,  <span class=\"hljs-comment\"># This is the default</span>\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Option 3: Use preprocessed HTML optimized for schema extraction</span>\n    fit_md_generator = DefaultMarkdownGenerator(\n        content_source=<span class=\"hljs-string\">\"fit_html\"</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n\n    <span class=\"hljs-comment\"># Use one of the generators in your crawler config</span>\n    config = CrawlerRunConfig(\n        markdown_generator=raw_md_generator  <span class=\"hljs-comment\"># Try each of the generators</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Markdown:\\n\"</span>, result.markdown.raw_markdown[:<span class=\"hljs-number\">500</span>])\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    <span class=\"hljs-keyword\">import</span> asyncio\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"html-source-options\">HTML Source Options</h3>\n<ul>\n<li>\n<p><strong><code>\"cleaned_html\"</code></strong> (default): Uses the HTML after it has been processed by the scraping strategy. This HTML is typically cleaner and more focused on content, with some boilerplate removed.</p>\n</li>\n<li>\n<p><strong><code>\"raw_html\"</code></strong>: Uses the original HTML directly from the webpage, before any cleaning or processing. This preserves more of the original content, but may include navigation bars, ads, footers, and other elements that might not be relevant to the main content.</p>\n</li>\n<li>\n<p><strong><code>\"fit_html\"</code></strong>: Uses HTML preprocessed for schema extraction. This HTML is optimized for structured data extraction and may have certain elements simplified or removed.</p>\n</li>\n</ul>\n<h3 id=\"when-to-use-each-option\">When to Use Each Option</h3>\n<ul>\n<li>Use <strong><code>\"cleaned_html\"</code></strong> (default) for most cases where you want a balance of content preservation and noise removal.</li>\n<li>Use <strong><code>\"raw_html\"</code></strong> when you need to preserve all original content, or when the cleaning process is removing content you actually want to keep.</li>\n<li>Use <strong><code>\"fit_html\"</code></strong> when working with structured data or when you need HTML that's optimized for schema extraction.</li>\n</ul>\n<hr>\n<h2 id=\"5-content-filters\">5. Content Filters</h2>\n<p><strong>Content filters</strong> selectively remove or rank sections of text before turning them into Markdown. This is especially helpful if your page has ads, nav bars, or other clutter you don’t want.</p>\n<h3 id=\"51-bm25contentfilter\">5.1 BM25ContentFilter</h3>\n<p>If you have a <strong>search query</strong>, BM25 is a good choice:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> CrawlerRunConfig\n\nbm25_filter = BM25ContentFilter(\n    user_query=<span class=\"hljs-string\">\"machine learning\"</span>,\n    bm25_threshold=<span class=\"hljs-number\">1.2</span>,\n    language=<span class=\"hljs-string\">\"english\"</span>\n)\n\nmd_generator = DefaultMarkdownGenerator(\n    content_filter=bm25_filter,\n    options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n)\n\nconfig = CrawlerRunConfig(markdown_generator=md_generator)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ul>\n<li><strong><code>user_query</code></strong>: The term you want to focus on. BM25 tries to keep only content blocks relevant to that query.  </li>\n<li><strong><code>bm25_threshold</code></strong>: Raise it to keep fewer blocks; lower it to keep more.  </li>\n<li><strong><code>use_stemming</code></strong> <em>(default <code>True</code>)</em>: Whether to apply stemming to the query and content.</li>\n<li><strong><code>language (str)</code></strong>: Language for stemming (default: 'english').</li>\n</ul>\n<p><strong>No query provided?</strong> BM25 tries to glean a context from page metadata, or you can simply treat it as a scorched-earth approach that discards text with low generic score. Realistically, you want to supply a query for best results.</p>\n<h3 id=\"52-pruningcontentfilter\">5.2 PruningContentFilter</h3>\n<p>If you <strong>don’t</strong> have a specific query, or if you just want a robust “junk remover,” use <code>PruningContentFilter</code>. It analyzes text density, link density, HTML structure, and known patterns (like “nav,” “footer”) to systematically prune extraneous or repetitive sections.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-cpp\">from crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n\nprune_filter = <span class=\"hljs-built_in\">PruningContentFilter</span>(\n    threshold=<span class=\"hljs-number\">0.5</span>,\n    threshold_type=<span class=\"hljs-string\">\"fixed\"</span>,  <span class=\"hljs-meta\"># or <span class=\"hljs-string\">\"dynamic\"</span></span>\n    min_word_threshold=<span class=\"hljs-number\">50</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ul>\n<li><strong><code>threshold</code></strong>: Score boundary. Blocks below this score get removed.  </li>\n<li><strong><code>threshold_type</code></strong>:  <ul>\n<li><code>\"fixed\"</code>: Straight comparison (<code>score &gt;= threshold</code> keeps the block).  </li>\n<li><code>\"dynamic\"</code>: The filter adjusts threshold in a data-driven manner.  </li>\n</ul>\n</li>\n<li><strong><code>min_word_threshold</code></strong>: Discard blocks under N words as likely too short or unhelpful.</li>\n</ul>\n<p><strong>When to Use PruningContentFilter</strong><br>\n- You want a broad cleanup without a user query.<br>\n- The page has lots of repeated sidebars, footers, or disclaimers that hamper text extraction.</p>\n<h3 id=\"53-llmcontentfilter\">5.3 LLMContentFilter</h3>\n<p>For intelligent content filtering and high-quality markdown generation, you can use the <strong>LLMContentFilter</strong>. This filter leverages LLMs to generate relevant markdown while preserving the original content's meaning and structure:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, LLMConfig, DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> LLMContentFilter\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Initialize LLM filter with specific instruction</span>\n    <span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-api-token\"</span>), <span class=\"hljs-comment\">#or use environment variable</span>\n        instruction=<span class=\"hljs-string\">\"\"\"\n        Focus on extracting the core educational content.\n        Include:\n        - Key concepts and explanations\n        - Important code examples\n        - Essential technical details\n        Exclude:\n        - Navigation elements\n        - Sidebars\n        - Footer content\n        Format the output as clean markdown with proper code blocks and headers.\n        \"\"\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">4096</span>,  <span class=\"hljs-comment\"># Adjust based on your needs</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n    md_generator = DefaultMarkdownGenerator(\n        content_filter=<span class=\"hljs-built_in\">filter</span>,\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n    config = CrawlerRunConfig(\n        markdown_generator=md_generator,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n        <span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown)  <span class=\"hljs-comment\"># Filtered markdown content</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Features:</strong>\n- <strong>Intelligent Filtering</strong>: Uses LLMs to understand and extract relevant content while maintaining context\n- <strong>Customizable Instructions</strong>: Tailor the filtering process with specific instructions\n- <strong>Chunk Processing</strong>: Handles large documents by processing them in chunks (controlled by <code>chunk_token_threshold</code>)\n- <strong>Parallel Processing</strong>: For better performance, use smaller <code>chunk_token_threshold</code> (e.g., 2048 or 4096) to enable parallel processing of content chunks</p>\n<p><strong>Two Common Use Cases:</strong></p>\n<ol>\n<li>\n<p><strong>Exact Content Preservation</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n    instruction=<span class=\"hljs-string\">\"\"\"\n    Extract the main educational content while preserving its original wording and substance completely.\n    1. Maintain the exact language and terminology\n    2. Keep all technical explanations and examples intact\n    3. Preserve the original flow and structure\n    4. Remove only clearly irrelevant elements like navigation menus and ads\n    \"\"\"</span>,\n    chunk_token_threshold=<span class=\"hljs-number\">4096</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Focused Content Extraction</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-built_in\">filter</span> = LLMContentFilter(\n    instruction=<span class=\"hljs-string\">\"\"\"\n    Focus on extracting specific types of content:\n    - Technical documentation\n    - Code examples\n    - API references\n    Reformat the content into clear, well-structured markdown\n    \"\"\"</span>,\n    chunk_token_threshold=<span class=\"hljs-number\">4096</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n</ol>\n<blockquote>\n<p><strong>Performance Tip</strong>: Set a smaller <code>chunk_token_threshold</code> (e.g., 2048 or 4096) to enable parallel processing of content chunks. The default value is infinity, which processes the entire content as a single chunk.</p>\n</blockquote>\n<hr>\n<h2 id=\"6-using-fit-markdown\">6. Using Fit Markdown</h2>\n<p>When a content filter is active, the library produces two forms of markdown inside <code>result.markdown</code>:</p>\n<p>1. <strong><code>raw_markdown</code></strong>: The full unfiltered markdown.<br>\n2. <strong><code>fit_markdown</code></strong>: A “fit” version where the filter has removed or trimmed noisy segments.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        markdown_generator=DefaultMarkdownGenerator(\n            content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.6</span>),\n            options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n        )\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://news.example.com/tech\"</span>, config=config)\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw markdown:\\n\"</span>, result.markdown)\n\n            <span class=\"hljs-comment\"># If a filter is used, we also have .fit_markdown:</span>\n            md_object = result.markdown  <span class=\"hljs-comment\"># or your equivalent</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Filtered markdown:\\n\"</span>, md_object.fit_markdown)\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"7-the-markdowngenerationresult-object\">7. The <code>MarkdownGenerationResult</code> Object</h2>\n<p>If your library stores detailed markdown output in an object like <code>MarkdownGenerationResult</code>, you’ll see fields such as:</p>\n<ul>\n<li><strong><code>raw_markdown</code></strong>: The direct HTML-to-markdown transformation (no filtering).  </li>\n<li><strong><code>markdown_with_citations</code></strong>: A version that moves links to reference-style footnotes.  </li>\n<li><strong><code>references_markdown</code></strong>: A separate string or section containing the gathered references.  </li>\n<li><strong><code>fit_markdown</code></strong>: The filtered markdown if you used a content filter.  </li>\n<li><strong><code>fit_html</code></strong>: The corresponding HTML snippet used to generate <code>fit_markdown</code> (helpful for debugging or advanced usage).</li>\n</ul>\n<p><strong>Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-swift\">md_obj <span class=\"hljs-operator\">=</span> result.markdown  # your library’s naming may vary\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"RAW:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.raw_markdown)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"CITED:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.markdown_with_citations)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"REFERENCES:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.references_markdown)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"FIT:<span class=\"hljs-subst\">\\n</span>\"</span>, md_obj.fit_markdown)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Why Does This Matter?</strong><br>\n- You can supply <code>raw_markdown</code> to an LLM if you want the entire text.<br>\n- Or feed <code>fit_markdown</code> into a vector database to reduce token usage.<br>\n- <code>references_markdown</code> can help you keep track of link provenance.</p>\n<hr>\n<p>Below is a <strong>revised section</strong> under “Combining Filters (BM25 + Pruning)” that demonstrates how you can run <strong>two</strong> passes of content filtering without re-crawling, by taking the HTML (or text) from a first pass and feeding it into the second filter. It uses real code patterns from the snippet you provided for <strong>BM25ContentFilter</strong>, which directly accepts <strong>HTML</strong> strings (and can also handle plain text with minimal adaptation).</p>\n<hr>\n<h2 id=\"8-combining-filters-bm25-pruning-in-two-passes\">8. Combining Filters (BM25 + Pruning) in Two Passes</h2>\n<p>You might want to <strong>prune out</strong> noisy boilerplate first (with <code>PruningContentFilter</code>), and then <strong>rank what’s left</strong> against a user query (with <code>BM25ContentFilter</code>). You don’t have to crawl the page twice. Instead:</p>\n<p>1. <strong>First pass</strong>: Apply <code>PruningContentFilter</code> directly to the raw HTML from <code>result.html</code> (the crawler’s downloaded HTML).<br>\n2. <strong>Second pass</strong>: Take the pruned HTML (or text) from step 1, and feed it into <code>BM25ContentFilter</code>, focusing on a user query.</p>\n<h3 id=\"two-pass-example\">Two-Pass Example</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter, BM25ContentFilter\n<span class=\"hljs-keyword\">from</span> bs4 <span class=\"hljs-keyword\">import</span> BeautifulSoup\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Crawl with minimal or no markdown generator, just get raw HTML</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># If you only want raw HTML, you can skip passing a markdown_generator</span>\n        <span class=\"hljs-comment\"># or provide one but focus on .html in this example</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/tech-article\"</span>, config=config)\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success <span class=\"hljs-keyword\">or</span> <span class=\"hljs-keyword\">not</span> result.html:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed or no HTML content.\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        raw_html = result.html\n\n        <span class=\"hljs-comment\"># 2. First pass: PruningContentFilter on raw HTML</span>\n        pruning_filter = PruningContentFilter(threshold=<span class=\"hljs-number\">0.5</span>, min_word_threshold=<span class=\"hljs-number\">50</span>)\n\n        <span class=\"hljs-comment\"># filter_content returns a list of \"text chunks\" or cleaned HTML sections</span>\n        pruned_chunks = pruning_filter.filter_content(raw_html)\n        <span class=\"hljs-comment\"># This list is basically pruned content blocks, presumably in HTML or text form</span>\n\n        <span class=\"hljs-comment\"># For demonstration, let's combine these chunks back into a single HTML-like string</span>\n        <span class=\"hljs-comment\"># or you could do further processing. It's up to your pipeline design.</span>\n        pruned_html = <span class=\"hljs-string\">\"\\n\"</span>.join(pruned_chunks)\n\n        <span class=\"hljs-comment\"># 3. Second pass: BM25ContentFilter with a user query</span>\n        bm25_filter = BM25ContentFilter(\n            user_query=<span class=\"hljs-string\">\"machine learning\"</span>,\n            bm25_threshold=<span class=\"hljs-number\">1.2</span>,\n            language=<span class=\"hljs-string\">\"english\"</span>\n        )\n\n        <span class=\"hljs-comment\"># returns a list of text chunks</span>\n        bm25_chunks = bm25_filter.filter_content(pruned_html)  \n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> bm25_chunks:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Nothing matched the BM25 query after pruning.\"</span>)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># 4. Combine or display final results</span>\n        final_text = <span class=\"hljs-string\">\"\\n---\\n\"</span>.join(bm25_chunks)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"==== PRUNED OUTPUT (first pass) ====\"</span>)\n        <span class=\"hljs-built_in\">print</span>(pruned_html[:<span class=\"hljs-number\">500</span>], <span class=\"hljs-string\">\"... (truncated)\"</span>)  <span class=\"hljs-comment\"># preview</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n==== BM25 OUTPUT (second pass) ====\"</span>)\n        <span class=\"hljs-built_in\">print</span>(final_text[:<span class=\"hljs-number\">500</span>], <span class=\"hljs-string\">\"... (truncated)\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"whats-happening\">What’s Happening?</h3>\n<p>1. <strong>Raw HTML</strong>: We crawl once and store the raw HTML in <code>result.html</code>.<br>\n2. <strong>PruningContentFilter</strong>: Takes HTML + optional parameters. It extracts blocks of text or partial HTML, removing headings/sections deemed “noise.” It returns a <strong>list of text chunks</strong>.<br>\n3. <strong>Combine or Transform</strong>: We join these pruned chunks back into a single HTML-like string. (Alternatively, you could store them in a list for further logic—whatever suits your pipeline.)<br>\n4. <strong>BM25ContentFilter</strong>: We feed the pruned string into <code>BM25ContentFilter</code> with a user query. This second pass further narrows the content to chunks relevant to “machine learning.”</p>\n<p><strong>No Re-Crawling</strong>: We used <code>raw_html</code> from the first pass, so there’s no need to run <code>arun()</code> again—<strong>no second network request</strong>.</p>\n<h3 id=\"tips-variations\">Tips &amp; Variations</h3>\n<ul>\n<li><strong>Plain Text vs. HTML</strong>: If your pruned output is mostly text, BM25 can still handle it; just keep in mind it expects a valid string input. If you supply partial HTML (like <code>\"&lt;p&gt;some text&lt;/p&gt;\"</code>), it will parse it as HTML.  </li>\n<li><strong>Chaining in a Single Pipeline</strong>: If your code supports it, you can chain multiple filters automatically. Otherwise, manual two-pass filtering (as shown) is straightforward.  </li>\n<li><strong>Adjust Thresholds</strong>: If you see too much or too little text in step one, tweak <code>threshold=0.5</code> or <code>min_word_threshold=50</code>. Similarly, <code>bm25_threshold=1.2</code> can be raised/lowered for more or fewer chunks in step two.</li>\n</ul>\n<h3 id=\"one-pass-combination\">One-Pass Combination?</h3>\n<p>If your codebase or pipeline design allows applying multiple filters in one pass, you could do so. But often it’s simpler—and more transparent—to run them sequentially, analyzing each step’s result.</p>\n<p><strong>Bottom Line</strong>: By <strong>manually chaining</strong> your filtering logic in two passes, you get powerful incremental control over the final content. First, remove “global” clutter with Pruning, then refine further with BM25-based query relevance—without incurring a second network crawl.</p>\n<hr>\n<h2 id=\"9-common-pitfalls-tips\">9. Common Pitfalls &amp; Tips</h2>\n<p>1. <strong>No Markdown Output?</strong><br>\n   - Make sure the crawler actually retrieved HTML. If the site is heavily JS-based, you may need to enable dynamic rendering or wait for elements.<br>\n   - Check if your content filter is too aggressive. Lower thresholds or disable the filter to see if content reappears.</p>\n<p>2. <strong>Performance Considerations</strong><br>\n   - Very large pages with multiple filters can be slower. Consider <code>cache_mode</code> to avoid re-downloading.<br>\n   - If your final use case is LLM ingestion, consider summarizing further or chunking big texts.</p>\n<p>3. <strong>Take Advantage of <code>fit_markdown</code></strong><br>\n   - Great for RAG pipelines, semantic search, or any scenario where extraneous boilerplate is unwanted.<br>\n   - Still verify the textual quality—some sites have crucial data in footers or sidebars.</p>\n<p>4. <strong>Adjusting <code>html2text</code> Options</strong><br>\n   - If you see lots of raw HTML slipping into the text, turn on <code>escape_html</code>.<br>\n   - If code blocks look messy, experiment with <code>mark_code</code> or <code>handle_code_in_pre</code>.</p>\n<hr>\n<h2 id=\"10-summary-next-steps\">10. Summary &amp; Next Steps</h2>\n<p>In this <strong>Markdown Generation Basics</strong> tutorial, you learned to:</p>\n<ul>\n<li>Configure the <strong>DefaultMarkdownGenerator</strong> with HTML-to-text options.  </li>\n<li>Select different HTML sources using the <code>content_source</code> parameter.  </li>\n<li>Use <strong>BM25ContentFilter</strong> for query-specific extraction or <strong>PruningContentFilter</strong> for general noise removal.  </li>\n<li>Distinguish between raw and filtered markdown (<code>fit_markdown</code>).  </li>\n<li>Leverage the <code>MarkdownGenerationResult</code> object to handle different forms of output (citations, references, etc.).</li>\n</ul>\n<p>Now you can produce high-quality Markdown from any website, focusing on exactly the content you need—an essential step for powering AI models, summarization pipelines, or knowledge-base queries.</p>\n<p><strong>Last Updated</strong>: 2025-01-01</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/page-interaction/",
    "title": "Page Interaction - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"page-interaction\">Page Interaction</h1>\n<p>Crawl4AI provides powerful features for interacting with <strong>dynamic</strong> webpages, handling JavaScript execution, waiting for conditions, and managing multi-step flows. By combining <strong>js_code</strong>, <strong>wait_for</strong>, and certain <strong>CrawlerRunConfig</strong> parameters, you can:</p>\n<ol>\n<li>Click “Load More” buttons  </li>\n<li>Fill forms and submit them  </li>\n<li>Wait for elements or data to appear  </li>\n<li>Reuse sessions across multiple steps  </li>\n</ol>\n<p>Below is a quick overview of how to do it.</p>\n<hr>\n<h2 id=\"1-javascript-execution\">1. JavaScript Execution</h2>\n<h3 id=\"basic-execution\">Basic Execution</h3>\n<p><strong><code>js_code</code></strong> in <strong><code>CrawlerRunConfig</code></strong> accepts either a single JS string or a list of JS snippets. It runs <strong>after</strong> <code>wait_for</code> and <code>delay_before_return_html</code> — so the page is fully loaded when your code executes.</p>\n<p><strong>Example</strong>: We'll scroll to the bottom of the page, then optionally click a \"Load More\" button.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Single JS command</span>\n    config = CrawlerRunConfig(\n        js_code=<span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># Example site</span>\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawled length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n    <span class=\"hljs-comment\"># Multiple commands</span>\n    js_commands = [\n        <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n        <span class=\"hljs-comment\"># 'More' link on Hacker News</span>\n        <span class=\"hljs-string\">\"document.querySelector('a.morelink')?.click();\"</span>,  \n    ]\n    config = CrawlerRunConfig(js_code=js_commands)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># Another pass</span>\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"After scroll+click, length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.cleaned_html))\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Relevant <code>CrawlerRunConfig</code> params</strong>:\n- <strong><code>js_code</code></strong>: JavaScript to run <strong>after</strong> <code>wait_for</code> and <code>delay_before_return_html</code> complete. Runs on the fully-loaded page.\n- <strong><code>js_code_before_wait</code></strong>: JavaScript to run <strong>before</strong> <code>wait_for</code>. Use when you need to trigger loading that <code>wait_for</code> then checks.\n- <strong><code>js_only</code></strong>: If set to <code>True</code> on subsequent calls, indicates we're continuing an existing session without a new full navigation.\n- <strong><code>session_id</code></strong>: If you want to keep the same page across multiple calls, specify an ID.</p>\n<h3 id=\"execution-order\">Execution Order</h3>\n<p>Understanding when your JavaScript runs relative to other pipeline steps:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-markdown\"><span class=\"hljs-bullet\">1.</span> Page navigation (page.goto)\n<span class=\"hljs-bullet\">2.</span> js<span class=\"hljs-emphasis\">_code_</span>before<span class=\"hljs-emphasis\">_wait     ← triggers loading / clicks tabs\n3. wait_</span>for                ← waits for content to appear\n<span class=\"hljs-bullet\">4.</span> delay<span class=\"hljs-emphasis\">_before_</span>return<span class=\"hljs-emphasis\">_html ← extra safety margin\n5. js_</span>code                 ← runs on the fully-loaded page\n<span class=\"hljs-bullet\">6.</span> flatten<span class=\"hljs-emphasis\">_shadow_</span>dom      ← if enabled\n<span class=\"hljs-bullet\">7.</span> page.content()          ← HTML capture\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>If you need JS to trigger something and then wait for the result, use <code>js_code_before_wait</code> + <code>wait_for</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    <span class=\"hljs-comment\"># Click a tab first</span>\n    js_code_before_wait=<span class=\"hljs-string\">\"document.querySelector('#specs-tab')?.click();\"</span>,\n    <span class=\"hljs-comment\"># Then wait for the tab content to appear</span>\n    wait_for=<span class=\"hljs-string\">\"css:#specs-panel .content\"</span>,\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"2-wait-conditions\">2. Wait Conditions</h2>\n<h3 id=\"21-css-based-waiting\">2.1 CSS-Based Waiting</h3>\n<p>Sometimes, you just want to wait for a specific element to appear. For example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Wait for at least 30 items on Hacker News</span>\n        wait_for=<span class=\"hljs-string\">\"css:.athing:nth-child(30)\"</span>  \n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"We have at least 30 items loaded!\"</span>)\n        <span class=\"hljs-comment\"># Rough check</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Total items in HTML:\"</span>, result.cleaned_html.count(<span class=\"hljs-string\">\"athing\"</span>))  \n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key param</strong>:\n- <strong><code>wait_for=\"css:...\"</code></strong>: Tells the crawler to wait until that CSS selector is present.</p>\n<h3 id=\"22-javascript-based-waiting\">2.2 JavaScript-Based Waiting</h3>\n<p>For more complex conditions (e.g., waiting for content length to exceed a threshold), prefix <code>js:</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">wait_condition</span> = <span class=\"hljs-string\">\"\"\"() =&gt; {\n    const items = document.querySelectorAll('.athing');\n    return items.length &gt; 50;  // Wait for at least 51 items\n}\"\"\"</span>\n\n<span class=\"hljs-attr\">config</span> = CrawlerRunConfig(wait_for=f<span class=\"hljs-string\">\"js:{wait_condition}\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Behind the Scenes</strong>: Crawl4AI keeps polling the JS function until it returns <code>true</code> or a timeout occurs.</p>\n<hr>\n<h2 id=\"3-handling-dynamic-content\">3. Handling Dynamic Content</h2>\n<p>Many modern sites require <strong>multiple steps</strong>: scrolling, clicking “Load More,” or updating via JavaScript. Below are typical patterns.</p>\n<h3 id=\"31-load-more-example-hacker-news-more-link\">3.1 Load More Example (Hacker News “More” Link)</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># Step 1: Load initial Hacker News page</span>\n    config = CrawlerRunConfig(\n        wait_for=<span class=\"hljs-string\">\"css:.athing:nth-child(30)\"</span>  <span class=\"hljs-comment\"># Wait for 30 items</span>\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n            config=config\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Initial items loaded.\"</span>)\n\n        <span class=\"hljs-comment\"># Step 2: Let's scroll and click the \"More\" link</span>\n        load_more_js = [\n            <span class=\"hljs-string\">\"window.scrollTo(0, document.body.scrollHeight);\"</span>,\n            <span class=\"hljs-comment\"># The \"More\" link at page bottom</span>\n            <span class=\"hljs-string\">\"document.querySelector('a.morelink')?.click();\"</span>  \n        ]\n\n        next_page_conf = CrawlerRunConfig(\n            js_code=load_more_js,\n            wait_for=<span class=\"hljs-string\">\"\"\"js:() =&gt; {\n                return document.querySelectorAll('.athing').length &gt; 30;\n            }\"\"\"</span>,\n            <span class=\"hljs-comment\"># Mark that we do not re-navigate, but run JS in the same session:</span>\n            js_only=<span class=\"hljs-literal\">True</span>,\n            session_id=<span class=\"hljs-string\">\"hn_session\"</span>\n        )\n\n        <span class=\"hljs-comment\"># Re-use the same crawler session</span>\n        result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,  <span class=\"hljs-comment\"># same URL but continuing session</span>\n            config=next_page_conf\n        )\n        total_items = result2.cleaned_html.count(<span class=\"hljs-string\">\"athing\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Items after load-more:\"</span>, total_items)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key params</strong>:\n- <strong><code>session_id=\"hn_session\"</code></strong>: Keep the same page across multiple calls to <code>arun()</code>.\n- <strong><code>js_only=True</code></strong>: We’re not performing a full reload, just applying JS in the existing page.\n- <strong><code>wait_for</code></strong> with <code>js:</code>: Wait for item count to grow beyond 30.</p>\n<hr>\n<h3 id=\"32-form-interaction\">3.2 Form Interaction</h3>\n<p>If the site has a search or login form, you can fill fields and submit them with <strong><code>js_code</code></strong>. For instance, if GitHub had a local search form:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">js_form_interaction = <span class=\"hljs-string\">\"\"\"\ndocument.querySelector('#your-search').value = 'TypeScript commits';\ndocument.querySelector('form').submit();\n\"\"\"</span>\n\nconfig = CrawlerRunConfig(\n    js_code=js_form_interaction,\n    wait_for=<span class=\"hljs-string\">\"css:.commit\"</span>\n)\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://github.com/search\"</span>, config=config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>In reality</strong>: Replace IDs or classes with the real site’s form selectors.</p>\n<hr>\n<h2 id=\"4-timing-control\">4. Timing Control</h2>\n<p>1. <strong><code>page_timeout</code></strong> (ms): Overall page load or script execution time limit.<br>\n2. <strong><code>delay_before_return_html</code></strong> (seconds): Wait an extra moment before capturing the final HTML.<br>\n3. <strong><code>mean_delay</code></strong> &amp; <strong><code>max_range</code></strong>: If you call <code>arun_many()</code> with multiple URLs, these add a random pause between each request.</p>\n<p><strong>Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    page_timeout=60000,  <span class=\"hljs-comment\"># 60s limit</span>\n    delay_before_return_html=2.5\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"5-multi-step-interaction-example\">5. Multi-Step Interaction Example</h2>\n<p>Below is a simplified script that does multiple “Load More” clicks on GitHub’s TypeScript commits page. It <strong>re-uses</strong> the same session to accumulate new commits each time. The code includes the relevant <strong><code>CrawlerRunConfig</code></strong> parameters you’d rely on.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">multi_page_commits</span>():\n    browser_cfg = BrowserConfig(\n        headless=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Visible for demonstration</span>\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n    session_id = <span class=\"hljs-string\">\"github_ts_commits\"</span>\n\n    base_wait = <span class=\"hljs-string\">\"\"\"js:() =&gt; {\n        const commits = document.querySelectorAll('li.Box-sc-g0xbh4-0 h4');\n        return commits.length &gt; 0;\n    }\"\"\"</span>\n\n    <span class=\"hljs-comment\"># Step 1: Load initial commits</span>\n    config1 = CrawlerRunConfig(\n        wait_for=base_wait,\n        session_id=session_id,\n        cache_mode=CacheMode.BYPASS,\n        <span class=\"hljs-comment\"># Not using js_only yet since it's our first load</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>,\n            config=config1\n        )\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Initial commits loaded. Count:\"</span>, result.cleaned_html.count(<span class=\"hljs-string\">\"commit\"</span>))\n\n        <span class=\"hljs-comment\"># Step 2: For subsequent pages, we run JS to click 'Next Page' if it exists</span>\n        js_next_page = <span class=\"hljs-string\">\"\"\"\n        const selector = 'a[data-testid=\"pagination-next-button\"]';\n        const button = document.querySelector(selector);\n        if (button) button.click();\n        \"\"\"</span>\n\n        <span class=\"hljs-comment\"># Wait until new commits appear</span>\n        wait_for_more = <span class=\"hljs-string\">\"\"\"js:() =&gt; {\n            const commits = document.querySelectorAll('li.Box-sc-g0xbh4-0 h4');\n            if (!window.firstCommit &amp;&amp; commits.length&gt;0) {\n                window.firstCommit = commits[0].textContent;\n                return false;\n            }\n            // If top commit changes, we have new commits\n            const topNow = commits[0]?.textContent.trim();\n            return topNow &amp;&amp; topNow !== window.firstCommit;\n        }\"\"\"</span>\n\n        <span class=\"hljs-keyword\">for</span> page <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">2</span>):  <span class=\"hljs-comment\"># let's do 2 more \"Next\" pages</span>\n            config_next = CrawlerRunConfig(\n                session_id=session_id,\n                js_code=js_next_page,\n                wait_for=wait_for_more,\n                js_only=<span class=\"hljs-literal\">True</span>,       <span class=\"hljs-comment\"># We're continuing from the open tab</span>\n                cache_mode=CacheMode.BYPASS\n            )\n            result2 = <span class=\"hljs-keyword\">await</span> crawler.arun(\n                url=<span class=\"hljs-string\">\"https://github.com/microsoft/TypeScript/commits/main\"</span>,\n                config=config_next\n            )\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Page <span class=\"hljs-subst\">{page+<span class=\"hljs-number\">2</span>}</span> commits count:\"</span>, result2.cleaned_html.count(<span class=\"hljs-string\">\"commit\"</span>))\n\n        <span class=\"hljs-comment\"># Optionally kill session</span>\n        <span class=\"hljs-keyword\">await</span> crawler.crawler_strategy.kill_session(session_id)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">await</span> multi_page_commits()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong>:</p>\n<ul>\n<li><strong><code>session_id</code></strong>: Keep the same page open.  </li>\n<li><strong><code>js_code</code></strong> + <strong><code>wait_for</code></strong> + <strong><code>js_only=True</code></strong>: We do partial refreshes, waiting for new commits to appear.  </li>\n<li><strong><code>cache_mode=CacheMode.BYPASS</code></strong> ensures we always see fresh data each step.</li>\n</ul>\n<hr>\n<h2 id=\"6-combine-interaction-with-extraction\">6. Combine Interaction with Extraction</h2>\n<p>Once dynamic content is loaded, you can attach an <strong><code>extraction_strategy</code></strong> (like <code>JsonCssExtractionStrategy</code> or <code>LLMExtractionStrategy</code>). For example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">from crawl4ai import JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Commits\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"li.Box-sc-g0xbh4-0\"</span>,\n    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h4.markdown-title\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span>\n    <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    session_id<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"ts_commits_session\"</span>,\n    js_code<span class=\"hljs-punctuation\">=</span>js_next_page,\n    wait_for<span class=\"hljs-punctuation\">=</span>wait_for_more,\n    extraction_strategy<span class=\"hljs-punctuation\">=</span>JsonCssExtractionStrategy<span class=\"hljs-punctuation\">(</span><span class=\"hljs-keyword\">schema</span><span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>When done, check <code>result.extracted_content</code> for the JSON.</p>\n<hr>\n<h2 id=\"7-shadow-dom-flattening\">7. Shadow DOM Flattening</h2>\n<p>Sites built with <strong>Web Components</strong> (Stencil, Lit, Shoelace, etc.) render content inside Shadow DOM — an encapsulated sub-tree that is invisible to normal page serialization. Set <code>flatten_shadow_dom=True</code> to extract it:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">config <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    flatten_shadow_dom<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    wait_until<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"load\"</span>,\n    delay_before_return_html<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">3.0</span>,  <span class=\"hljs-comment\"># give components time to hydrate</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This walks all shadow trees, resolves <code>&lt;slot&gt;</code> projections, and produces flat HTML. It also force-opens closed shadow roots via an init script. For details and a full example, see <a href=\"../content-selection/#31-flattening-shadow-dom\">Flattening Shadow DOM</a> and <a href=\"https://github.com/unclecode/crawl4ai/blob/main/docs/examples/shadow_dom_crawling.py\"><code>shadow_dom_crawling.py</code></a>.</p>\n<hr>\n<h2 id=\"8-relevant-crawlerrunconfig-parameters\">8. Relevant <code>CrawlerRunConfig</code> Parameters</h2>\n<p>Below are the key interaction-related parameters in <code>CrawlerRunConfig</code>. For a full list, see <a href=\"../../api/parameters/\">Configuration Parameters</a>.</p>\n<ul>\n<li><strong><code>js_code</code></strong>: JavaScript to run after <code>wait_for</code> + <code>delay_before_return_html</code>, on the fully-loaded page.</li>\n<li><strong><code>js_code_before_wait</code></strong>: JavaScript to run before <code>wait_for</code>. For triggering loading that <code>wait_for</code> then checks.</li>\n<li><strong><code>js_only</code></strong>: If <code>True</code>, no new page navigation—only JS in the existing session.</li>\n<li><strong><code>wait_for</code></strong>: CSS (<code>\"css:...\"</code>) or JS (<code>\"js:...\"</code>) expression to wait for.</li>\n<li><strong><code>session_id</code></strong>: Reuse the same page across calls.</li>\n<li><strong><code>cache_mode</code></strong>: Whether to read/write from the cache or bypass.</li>\n<li><strong><code>flatten_shadow_dom</code></strong>: Flatten Shadow DOM content into the light DOM before capture.</li>\n<li><strong><code>process_iframes</code></strong>: Inline iframe content into the main document.</li>\n<li><strong><code>remove_overlay_elements</code></strong>: Remove certain popups automatically.</li>\n<li><strong><code>remove_consent_popups</code></strong>: Remove GDPR/cookie consent popups from known CMP providers (OneTrust, Cookiebot, Didomi, etc.).</li>\n<li><strong><code>simulate_user</code>, <code>override_navigator</code>, <code>magic</code></strong>: Anti-bot or \"human-like\" interactions.</li>\n</ul>\n<hr>\n<h2 id=\"9-conclusion\">9. Conclusion</h2>\n<p>Crawl4AI's <strong>page interaction</strong> features let you:</p>\n<p>1. <strong>Execute JavaScript</strong> for scrolling, clicks, or form filling.<br>\n2. <strong>Wait</strong> for CSS or custom JS conditions before capturing data.<br>\n3. <strong>Handle</strong> multi-step flows (like “Load More”) with partial reloads or persistent sessions.<br>\n4. <strong>Flatten Shadow DOM</strong> on Web Component sites to extract hidden content.\n5. Combine with <strong>structured extraction</strong> for dynamic sites.</p>\n<p>With these tools, you can scrape modern, interactive webpages confidently. For advanced hooking, user simulation, or in-depth config, check the <a href=\"../../api/parameters/\">API reference</a> or related advanced docs. Happy scripting!</p>\n<hr>\n<h2 id=\"10-virtual-scrolling\">10. Virtual Scrolling</h2>\n<p>For sites that use <strong>virtual scrolling</strong> (where content is replaced rather than appended as you scroll, like Twitter or Instagram), Crawl4AI provides a dedicated <code>VirtualScrollConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, VirtualScrollConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">crawl_twitter_timeline</span>():\n    <span class=\"hljs-comment\"># Configure virtual scroll for Twitter-like feeds</span>\n    virtual_config = VirtualScrollConfig(\n        container_selector=<span class=\"hljs-string\">\"[data-testid='primaryColumn']\"</span>,  <span class=\"hljs-comment\"># Twitter's main column</span>\n        scroll_count=<span class=\"hljs-number\">30</span>,                <span class=\"hljs-comment\"># Scroll 30 times</span>\n        scroll_by=<span class=\"hljs-string\">\"container_height\"</span>,   <span class=\"hljs-comment\"># Scroll by container height each time</span>\n        wait_after_scroll=<span class=\"hljs-number\">1.0</span>          <span class=\"hljs-comment\"># Wait 1 second after each scroll</span>\n    )\n\n    config = CrawlerRunConfig(\n        virtual_scroll_config=virtual_config\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://twitter.com/search?q=AI\"</span>,\n            config=config\n        )\n        <span class=\"hljs-comment\"># result.html now contains ALL tweets from the virtual scroll</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"virtual-scroll-vs-javascript-scrolling\">Virtual Scroll vs JavaScript Scrolling</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Feature</th>\n<th>Virtual Scroll</th>\n<th>JS Code Scrolling</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Use Case</strong></td>\n<td>Content replaced during scroll</td>\n<td>Content appended or simple scroll</td>\n</tr>\n<tr>\n<td><strong>Configuration</strong></td>\n<td><code>VirtualScrollConfig</code> object</td>\n<td><code>js_code</code> with scroll commands</td>\n</tr>\n<tr>\n<td><strong>Automatic Merging</strong></td>\n<td>Yes - merges all unique content</td>\n<td>No - captures final state only</td>\n</tr>\n<tr>\n<td><strong>Best For</strong></td>\n<td>Twitter, Instagram, virtual tables</td>\n<td>Traditional pages, load more buttons</td>\n</tr>\n</tbody>\n</table>\n<p>For detailed examples and configuration options, see the <a href=\"../../advanced/virtual-scroll/\">Virtual Scroll documentation</a>.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/quickstart/",
    "title": "Quick Start - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"getting-started-with-crawl4ai\">Getting Started with Crawl4AI</h1>\n<p>Welcome to <strong>Crawl4AI</strong>, an open-source LLM-friendly Web Crawler &amp; Scraper. In this tutorial, you’ll:</p>\n<ol>\n<li>Run your <strong>first crawl</strong> using minimal configuration.  </li>\n<li>Generate <strong>Markdown</strong> output (and learn how it’s influenced by content filters).  </li>\n<li>Experiment with a simple <strong>CSS-based extraction</strong> strategy.  </li>\n<li>See a glimpse of <strong>LLM-based extraction</strong> (including open-source and closed-source model options).  </li>\n<li>Crawl a <strong>dynamic</strong> page that loads content via JavaScript.</li>\n</ol>\n<hr>\n<h2 id=\"1-introduction\">1. Introduction</h2>\n<p>Crawl4AI provides:</p>\n<ul>\n<li>An asynchronous crawler, <strong><code>AsyncWebCrawler</code></strong>.  </li>\n<li>Configurable browser and run settings via <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong>.  </li>\n<li>Automatic HTML-to-Markdown conversion via <strong><code>DefaultMarkdownGenerator</code></strong> (supports optional filters).  </li>\n<li>Multiple extraction strategies (LLM-based or “traditional” CSS/XPath-based).</li>\n</ul>\n<p>By the end of this guide, you’ll have performed a basic crawl, generated Markdown, tried out two extraction strategies, and crawled a dynamic page that uses “Load More” buttons or JavaScript updates.</p>\n<hr>\n<h2 id=\"2-your-first-crawl\">2. Your First Crawl</h2>\n<p>Here’s a minimal Python script that creates an <strong><code>AsyncWebCrawler</code></strong>, fetches a webpage, and prints the first 300 characters of its Markdown output:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>)\n        <span class=\"hljs-built_in\">print</span>(result.markdown[:<span class=\"hljs-number\">300</span>])  <span class=\"hljs-comment\"># Print first 300 chars</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What’s happening?</strong>\n- <strong><code>AsyncWebCrawler</code></strong> launches a headless browser (Chromium by default).\n- It fetches <code>https://example.com</code>.\n- Crawl4AI automatically converts the HTML into Markdown.</p>\n<p>You now have a simple, working crawl!</p>\n<hr>\n<h2 id=\"3-basic-configuration-light-introduction\">3. Basic Configuration (Light Introduction)</h2>\n<p>Crawl4AI’s crawler can be heavily customized using two main classes:</p>\n<p>1. <strong><code>BrowserConfig</code></strong>: Controls browser behavior (headless or full UI, user agent, JavaScript toggles, etc.).<br>\n2. <strong><code>CrawlerRunConfig</code></strong>: Controls how each crawl runs (caching, extraction, timeouts, hooking, etc.).</p>\n<p>Below is an example with minimal usage:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_conf = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)  <span class=\"hljs-comment\"># or False to see the browser</span>\n    run_conf = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_conf) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_conf\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<blockquote>\n<p>IMPORTANT: By default cache mode is set to <code>CacheMode.BYPASS</code> to have fresh content. Set <code>CacheMode.ENABLED</code> to enable caching.</p>\n</blockquote>\n<p>We’ll explore more advanced config in later tutorials (like enabling proxies, PDF output, multi-tab sessions, etc.). For now, just note how you pass these objects to manage crawling.</p>\n<hr>\n<h2 id=\"4-generating-markdown-output\">4. Generating Markdown Output</h2>\n<p>By default, Crawl4AI automatically generates Markdown from each crawled page. However, the exact output depends on whether you specify a <strong>markdown generator</strong> or <strong>content filter</strong>.</p>\n<ul>\n<li><strong><code>result.markdown</code></strong>:<br>\n  The direct HTML-to-Markdown conversion.  </li>\n<li><strong><code>result.markdown.fit_markdown</code></strong>:<br>\n  The same content after applying any configured <strong>content filter</strong> (e.g., <code>PruningContentFilter</code>).</li>\n</ul>\n<h3 id=\"example-using-a-filter-with-defaultmarkdowngenerator\">Example: Using a Filter with <code>DefaultMarkdownGenerator</code></h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai.content_filter_strategy <span class=\"hljs-keyword\">import</span> PruningContentFilter\n<span class=\"hljs-keyword\">from</span> crawl4ai.markdown_generation_strategy <span class=\"hljs-keyword\">import</span> DefaultMarkdownGenerator\n\nmd_generator = DefaultMarkdownGenerator(\n    content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.4</span>, threshold_type=<span class=\"hljs-string\">\"fixed\"</span>)\n)\n\nconfig = CrawlerRunConfig(\n    cache_mode=CacheMode.BYPASS,\n    markdown_generator=md_generator\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>, config=config)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Raw Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown))\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Fit Markdown length:\"</span>, <span class=\"hljs-built_in\">len</span>(result.markdown.fit_markdown))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Note</strong>: If you do <strong>not</strong> specify a content filter or markdown generator, you’ll typically see only the raw Markdown. <code>PruningContentFilter</code> may adds around <code>50ms</code> in processing time. We’ll dive deeper into these strategies in a dedicated <strong>Markdown Generation</strong> tutorial.</p>\n<hr>\n<h2 id=\"5-simple-data-extraction-css-based\">5. Simple Data Extraction (CSS-based)</h2>\n<p>Crawl4AI can also extract structured data (JSON) using CSS or XPath selectors. Below is a minimal CSS-based example:</p>\n<blockquote>\n<p><strong>New!</strong> Crawl4AI now provides a powerful utility to automatically generate extraction schemas using LLM. This is a one-time cost that gives you a reusable schema for fast, LLM-free extractions:</p>\n</blockquote>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Generate a schema (one-time cost)</span>\nhtml = <span class=\"hljs-string\">\"&lt;div class='product'&gt;&lt;h2&gt;Gaming Laptop&lt;/h2&gt;&lt;span class='price'&gt;$999.99&lt;/span&gt;&lt;/div&gt;\"</span>\n\n<span class=\"hljs-comment\"># Using OpenAI (requires API token)</span>\nschema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-openai-token\"</span>)  <span class=\"hljs-comment\"># Required for OpenAI</span>\n)\n\n<span class=\"hljs-comment\"># Or using Ollama (open source, no token needed)</span>\nschema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama3.3\"</span>, api_token=<span class=\"hljs-literal\">None</span>)  <span class=\"hljs-comment\"># Not needed for Ollama</span>\n)\n\n<span class=\"hljs-comment\"># Use the schema for fast, repeated extractions</span>\nstrategy = JsonCssExtractionStrategy(schema)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>For a complete guide on schema generation and advanced usage, see <a href=\"../../extraction/no-llm-strategies/\">No-LLM Extraction Strategies</a>.</p>\n<p>Here's a basic extraction example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Example Items\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.item\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>},\n            {<span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"link\"</span>, <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"a\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"href\"</span>}\n        ]\n    }\n\n    raw_html = <span class=\"hljs-string\">\"&lt;div class='item'&gt;&lt;h2&gt;Item 1&lt;/h2&gt;&lt;a href='https://example.com/item1'&gt;Link 1&lt;/a&gt;&lt;/div&gt;\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"raw://\"</span> + raw_html,\n            config=CrawlerRunConfig(\n                cache_mode=CacheMode.BYPASS,\n                extraction_strategy=JsonCssExtractionStrategy(schema)\n            )\n        )\n        <span class=\"hljs-comment\"># The JSON output is stored in 'extracted_content'</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(data)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Why is this helpful?</strong>\n- Great for repetitive page structures (e.g., item listings, articles).\n- No AI usage or costs.\n- The crawler returns a JSON string you can parse or store.</p>\n<blockquote>\n<p>Tips: You can pass raw HTML to the crawler instead of a URL. To do so, prefix the HTML with <code>raw://</code>.</p>\n</blockquote>\n<hr>\n<h2 id=\"6-simple-data-extraction-llm-based\">6. Simple Data Extraction (LLM-based)</h2>\n<p>For more complex or irregular pages, a language model can parse text intelligently into a structure you define. Crawl4AI supports <strong>open-source</strong> or <strong>closed-source</strong> providers:</p>\n<ul>\n<li><strong>Open-Source Models</strong> (e.g., <code>ollama/llama3.3</code>, <code>no_token</code>)  </li>\n<li><strong>OpenAI Models</strong> (e.g., <code>openai/gpt-4</code>, requires <code>api_token</code>)  </li>\n<li>Or any provider supported by the underlying library</li>\n</ul>\n<p>Below is an example using <strong>open-source</strong> style (no token) and closed-source:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">OpenAIModelFee</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    model_name: <span class=\"hljs-built_in\">str</span> = Field(..., description=<span class=\"hljs-string\">\"Name of the OpenAI model.\"</span>)\n    input_fee: <span class=\"hljs-built_in\">str</span> = Field(..., description=<span class=\"hljs-string\">\"Fee for input token for the OpenAI model.\"</span>)\n    output_fee: <span class=\"hljs-built_in\">str</span> = Field(\n        ..., description=<span class=\"hljs-string\">\"Fee for output token for the OpenAI model.\"</span>\n    )\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_structured_data_using_llm</span>(<span class=\"hljs-params\">\n    provider: <span class=\"hljs-built_in\">str</span>, api_token: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>, extra_headers: <span class=\"hljs-type\">Dict</span>[<span class=\"hljs-built_in\">str</span>, <span class=\"hljs-built_in\">str</span>] = <span class=\"hljs-literal\">None</span>\n</span>):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n--- Extracting Structured Data with <span class=\"hljs-subst\">{provider}</span> ---\"</span>)\n\n    <span class=\"hljs-keyword\">if</span> api_token <span class=\"hljs-keyword\">is</span> <span class=\"hljs-literal\">None</span> <span class=\"hljs-keyword\">and</span> provider != <span class=\"hljs-string\">\"ollama\"</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"API token is required for <span class=\"hljs-subst\">{provider}</span>. Skipping this example.\"</span>)\n        <span class=\"hljs-keyword\">return</span>\n\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)\n\n    extra_args = {<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0</span>, <span class=\"hljs-string\">\"top_p\"</span>: <span class=\"hljs-number\">0.9</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">2000</span>}\n    <span class=\"hljs-keyword\">if</span> extra_headers:\n        extra_args[<span class=\"hljs-string\">\"extra_headers\"</span>] = extra_headers\n\n    crawler_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        word_count_threshold=<span class=\"hljs-number\">1</span>,\n        page_timeout=<span class=\"hljs-number\">80000</span>,\n        extraction_strategy=LLMExtractionStrategy(\n            llm_config = LLMConfig(provider=provider,api_token=api_token),\n            schema=OpenAIModelFee.model_json_schema(),\n            extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n            instruction=<span class=\"hljs-string\">\"\"\"From the crawled content, extract all mentioned model names along with their fees for input and output tokens. \n            Do not miss any models in the entire content.\"\"\"</span>,\n            extra_args=extra_args,\n        ),\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://openai.com/api/pricing/\"</span>, config=crawler_config\n        )\n        <span class=\"hljs-built_in\">print</span>(result.extracted_content)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n\n    asyncio.run(\n        extract_structured_data_using_llm(\n            provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>, api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>)\n        )\n    )\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What’s happening?</strong>\n- We define a Pydantic schema (<code>PricingInfo</code>) describing the fields we want.\n- The LLM extraction strategy uses that schema and your instructions to transform raw text into structured JSON.\n- Depending on the <strong>provider</strong> and <strong>api_token</strong>, you can use local models or a remote API.</p>\n<hr>\n<h2 id=\"7-adaptive-crawling-new\">7. Adaptive Crawling (New!)</h2>\n<p>Crawl4AI now includes intelligent adaptive crawling that automatically determines when sufficient information has been gathered. Here's a quick example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, AdaptiveCrawler\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">adaptive_example</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        adaptive = AdaptiveCrawler(crawler)\n\n        <span class=\"hljs-comment\"># Start adaptive crawling</span>\n        result = <span class=\"hljs-keyword\">await</span> adaptive.digest(\n            start_url=<span class=\"hljs-string\">\"https://docs.python.org/3/\"</span>,\n            query=<span class=\"hljs-string\">\"async context managers\"</span>\n        )\n\n        <span class=\"hljs-comment\"># View results</span>\n        adaptive.print_stats()\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.crawled_urls)}</span> pages\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Achieved <span class=\"hljs-subst\">{adaptive.confidence:<span class=\"hljs-number\">.0</span>%}</span> confidence\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(adaptive_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What's special about adaptive crawling?</strong>\n- <strong>Automatic stopping</strong>: Stops when sufficient information is gathered\n- <strong>Intelligent link selection</strong>: Follows only relevant links\n- <strong>Confidence scoring</strong>: Know how complete your information is</p>\n<p><a href=\"../adaptive-crawling/\">Learn more about Adaptive Crawling →</a></p>\n<hr>\n<h2 id=\"8-multi-url-concurrency-preview\">8. Multi-URL Concurrency (Preview)</h2>\n<p>If you need to crawl multiple URLs in <strong>parallel</strong>, you can use <code>arun_many()</code>. By default, Crawl4AI employs a <strong>MemoryAdaptiveDispatcher</strong>, automatically adjusting concurrency based on system resources. Here’s a quick glimpse:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">quick_parallel_example</span>():\n    urls = [\n        <span class=\"hljs-string\">\"https://example.com/page1\"</span>,\n        <span class=\"hljs-string\">\"https://example.com/page2\"</span>,\n        <span class=\"hljs-string\">\"https://example.com/page3\"</span>\n    ]\n\n    run_conf = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        stream=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Enable streaming mode</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Stream results as they complete</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls, config=run_conf):\n            <span class=\"hljs-keyword\">if</span> result.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] <span class=\"hljs-subst\">{result.url}</span>, length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result.markdown.raw_markdown)}</span>\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[ERROR] <span class=\"hljs-subst\">{result.url}</span> =&gt; <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n        <span class=\"hljs-comment\"># Or get all results at once (default behavior)</span>\n        run_conf = run_conf.clone(stream=<span class=\"hljs-literal\">False</span>)\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls, config=run_conf)\n        <span class=\"hljs-keyword\">for</span> res <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-keyword\">if</span> res.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[OK] <span class=\"hljs-subst\">{res.url}</span>, length: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(res.markdown.raw_markdown)}</span>\"</span>)\n            <span class=\"hljs-keyword\">else</span>:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"[ERROR] <span class=\"hljs-subst\">{res.url}</span> =&gt; <span class=\"hljs-subst\">{res.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(quick_parallel_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The example above shows two ways to handle multiple URLs:\n1. <strong>Streaming mode</strong> (<code>stream=True</code>): Process results as they become available using <code>async for</code>\n2. <strong>Batch mode</strong> (<code>stream=False</code>): Wait for all results to complete</p>\n<p>For more advanced concurrency (e.g., a <strong>semaphore-based</strong> approach, <strong>adaptive memory usage throttling</strong>, or customized rate limiting), see <a href=\"../../advanced/multi-url-crawling/\">Advanced Multi-URL Crawling</a>.</p>\n<hr>\n<h2 id=\"8-dynamic-content-example\">8. Dynamic Content Example</h2>\n<p>Some sites require multiple “page clicks” or dynamic JavaScript updates. Below is an example showing how to <strong>click</strong> a “Next Page” button and wait for new commits to load on GitHub, using <strong><code>BrowserConfig</code></strong> and <strong><code>CrawlerRunConfig</code></strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_structured_data_using_css_extractor</span>():\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n--- Using JsonCssExtractionStrategy for Fast Structured Output ---\"</span>)\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"KidoCode Courses\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"section.charge-methodology .w-tab-content &gt; div\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"section_title\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h3.heading-50\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"section_description\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".charge-content\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".text-block-93\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_description\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".course-content-text\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>,\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"course_icon\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".image-92\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"attribute\"</span>,\n                <span class=\"hljs-string\">\"attribute\"</span>: <span class=\"hljs-string\">\"src\"</span>,\n            },\n        ],\n    }\n\n    browser_config = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>, java_script_enabled=<span class=\"hljs-literal\">True</span>)\n\n    js_click_tabs = <span class=\"hljs-string\">\"\"\"\n    (async () =&gt; {\n        const tabs = document.querySelectorAll(\"section.charge-methodology .tabs-menu-3 &gt; div\");\n        for(let tab of tabs) {\n            tab.scrollIntoView();\n            tab.click();\n            await new Promise(r =&gt; setTimeout(r, 500));\n        }\n    })();\n    \"\"\"</span>\n\n    crawler_config = CrawlerRunConfig(\n        cache_mode=CacheMode.BYPASS,\n        extraction_strategy=JsonCssExtractionStrategy(schema),\n        js_code=[js_click_tabs],\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.kidocode.com/degrees/technology\"</span>, config=crawler_config\n        )\n\n        companies = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Successfully extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(companies)}</span> companies\"</span>)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(companies[<span class=\"hljs-number\">0</span>], indent=<span class=\"hljs-number\">2</span>))\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">await</span> extract_structured_data_using_css_extractor()\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong>:</p>\n<ul>\n<li><strong><code>BrowserConfig(headless=False)</code></strong>: We want to watch it click “Next Page.”  </li>\n<li><strong><code>CrawlerRunConfig(...)</code></strong>: We specify the extraction strategy, pass <code>session_id</code> to reuse the same page.  </li>\n<li><strong><code>js_code</code></strong> and <strong><code>wait_for</code></strong> are used for subsequent pages (<code>page &gt; 0</code>) to click the “Next” button and wait for new commits to load.  </li>\n<li><strong><code>js_only=True</code></strong> indicates we’re not re-navigating but continuing the existing session.  </li>\n<li>Finally, we call <code>kill_session()</code> to clean up the page and browser session.</li>\n</ul>\n<hr>\n<h2 id=\"9-next-steps\">9. Next Steps</h2>\n<p>Congratulations! You have:</p>\n<ol>\n<li>Performed a basic crawl and printed Markdown.  </li>\n<li>Used <strong>content filters</strong> with a markdown generator.  </li>\n<li>Extracted JSON via <strong>CSS</strong> or <strong>LLM</strong> strategies.  </li>\n<li>Handled <strong>dynamic</strong> pages with JavaScript triggers.</li>\n</ol>\n<p>If you’re ready for more, check out:</p>\n<ul>\n<li><strong>Installation</strong>: A deeper dive into advanced installs, Docker usage (experimental), or optional dependencies.  </li>\n<li><strong>Hooks &amp; Auth</strong>: Learn how to run custom JavaScript or handle logins with cookies, local storage, etc.  </li>\n<li><strong>Deployment</strong>: Explore ephemeral testing in Docker or plan for the upcoming stable Docker release.  </li>\n<li><strong>Browser Management</strong>: Delve into user simulation, stealth modes, and concurrency best practices.  </li>\n</ul>\n<p>Crawl4AI is a powerful, flexible tool. Enjoy building out your scrapers, data pipelines, or AI-driven extraction flows. Happy crawling!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/simple-crawling/",
    "title": "Simple Crawling - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"simple-crawling\">Simple Crawling</h1>\n<p>This guide covers the basics of web crawling with Crawl4AI. You'll learn how to set up a crawler, make your first request, and understand the response.</p>\n<h2 id=\"basic-usage\">Basic Usage</h2>\n<p>Set up a simple crawl using <code>BrowserConfig</code> and <code>CrawlerRunConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig()  <span class=\"hljs-comment\"># Default browser configuration</span>\n    run_config = CrawlerRunConfig()   <span class=\"hljs-comment\"># Default crawl run configuration</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n        <span class=\"hljs-built_in\">print</span>(result.markdown)  <span class=\"hljs-comment\"># Print clean markdown content</span>\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"understanding-the-response\">Understanding the Response</h2>\n<p>The <code>arun()</code> method returns a <code>CrawlResult</code> object with several useful properties. Here's a quick overview (see <a href=\"../../api/crawl-result/\">CrawlResult</a> for complete details):</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = CrawlerRunConfig(\n    markdown_generator=DefaultMarkdownGenerator(\n        content_filter=PruningContentFilter(threshold=<span class=\"hljs-number\">0.6</span>),\n        options={<span class=\"hljs-string\">\"ignore_links\"</span>: <span class=\"hljs-literal\">True</span>}\n    )\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n    config=config\n)\n\n<span class=\"hljs-comment\"># Different content formats</span>\n<span class=\"hljs-built_in\">print</span>(result.html)         <span class=\"hljs-comment\"># Raw HTML</span>\n<span class=\"hljs-built_in\">print</span>(result.cleaned_html) <span class=\"hljs-comment\"># Cleaned HTML</span>\n<span class=\"hljs-built_in\">print</span>(result.markdown.raw_markdown) <span class=\"hljs-comment\"># Raw markdown from cleaned html</span>\n<span class=\"hljs-built_in\">print</span>(result.markdown.fit_markdown) <span class=\"hljs-comment\"># Most relevant content in markdown</span>\n\n<span class=\"hljs-comment\"># Check success status</span>\n<span class=\"hljs-built_in\">print</span>(result.success)      <span class=\"hljs-comment\"># True if crawl succeeded</span>\n<span class=\"hljs-built_in\">print</span>(result.status_code)  <span class=\"hljs-comment\"># HTTP status code (e.g., 200, 404)</span>\n\n<span class=\"hljs-comment\"># Access extracted media and links</span>\n<span class=\"hljs-built_in\">print</span>(result.media)        <span class=\"hljs-comment\"># Dictionary of found media (images, videos, audio)</span>\n<span class=\"hljs-built_in\">print</span>(result.links)        <span class=\"hljs-comment\"># Dictionary of internal and external links</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"adding-basic-options\">Adding Basic Options</h2>\n<p>Customize your crawl using <code>CrawlerRunConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">run_config = CrawlerRunConfig(\n    word_count_threshold=<span class=\"hljs-number\">10</span>,        <span class=\"hljs-comment\"># Minimum words per content block</span>\n    exclude_external_links=<span class=\"hljs-literal\">True</span>,    <span class=\"hljs-comment\"># Remove external links</span>\n    remove_overlay_elements=<span class=\"hljs-literal\">True</span>,   <span class=\"hljs-comment\"># Remove popups/modals</span>\n    process_iframes=<span class=\"hljs-literal\">True</span>           <span class=\"hljs-comment\"># Process iframe content</span>\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n    config=run_config\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"handling-errors\">Handling Errors</h2>\n<p>Always check if the crawl was successful:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">run_config = CrawlerRunConfig()\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n\n<span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Status code: <span class=\"hljs-subst\">{result.status_code}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"logging-and-debugging\">Logging and Debugging</h2>\n<p>Enable verbose logging in <code>BrowserConfig</code>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\">browser_config = BrowserConfig(verbose=True)\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>(<span class=\"hljs-params\">config=browser_config</span>) <span class=\"hljs-keyword\">as</span> crawler:\n    run_config</span> = CrawlerRunConfig()\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=<span class=\"hljs-string\">\"https://example.com\"</span>, config=run_config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"complete-example\">Complete Example</h2>\n<p>Here's a more comprehensive example demonstrating common usage patterns:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler\n<span class=\"hljs-keyword\">from</span> crawl4ai.async_configs <span class=\"hljs-keyword\">import</span> BrowserConfig, CrawlerRunConfig, CacheMode\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    browser_config = BrowserConfig(verbose=<span class=\"hljs-literal\">True</span>)\n    run_config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># Content filtering</span>\n        word_count_threshold=<span class=\"hljs-number\">10</span>,\n        excluded_tags=[<span class=\"hljs-string\">'form'</span>, <span class=\"hljs-string\">'header'</span>],\n        exclude_external_links=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Content processing</span>\n        process_iframes=<span class=\"hljs-literal\">True</span>,\n        remove_overlay_elements=<span class=\"hljs-literal\">True</span>,\n\n        <span class=\"hljs-comment\"># Cache control</span>\n        cache_mode=CacheMode.ENABLED  <span class=\"hljs-comment\"># Use cache if available</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_config) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=run_config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># Print clean content</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Content:\"</span>, result.markdown[:<span class=\"hljs-number\">500</span>])  <span class=\"hljs-comment\"># First 500 chars</span>\n\n            <span class=\"hljs-comment\"># Process images</span>\n            <span class=\"hljs-keyword\">for</span> image <span class=\"hljs-keyword\">in</span> result.media[<span class=\"hljs-string\">\"images\"</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found image: <span class=\"hljs-subst\">{image[<span class=\"hljs-string\">'src'</span>]}</span>\"</span>)\n\n            <span class=\"hljs-comment\"># Process links</span>\n            <span class=\"hljs-keyword\">for</span> link <span class=\"hljs-keyword\">in</span> result.links[<span class=\"hljs-string\">\"internal\"</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Internal link: <span class=\"hljs-subst\">{link[<span class=\"hljs-string\">'href'</span>]}</span>\"</span>)\n\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/table_extraction/",
    "title": "Table Extraction Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"table-extraction-strategies\">Table Extraction Strategies</h1>\n<h2 id=\"overview\">Overview</h2>\n<p><strong>New in v0.7.3+</strong>: Table extraction now follows the <strong>Strategy Design Pattern</strong>, providing unprecedented flexibility and power for handling different table structures. Don't worry - <strong>your existing code still works!</strong> We maintain full backward compatibility while offering new capabilities.</p>\n<h3 id=\"whats-changed\">What's Changed?</h3>\n<ul>\n<li><strong>Architecture</strong>: Table extraction now uses pluggable strategies</li>\n<li><strong>Backward Compatible</strong>: Your existing code with <code>table_score_threshold</code> continues to work</li>\n<li><strong>More Power</strong>: Choose from multiple strategies or create your own</li>\n<li><strong>Same Default Behavior</strong>: By default, uses <code>DefaultTableExtraction</code> (same as before)</li>\n</ul>\n<h3 id=\"key-points\">Key Points</h3>\n<p>✅ <strong>Old code still works</strong> - No breaking changes<br>\n✅ <strong>Same default behavior</strong> - Uses the proven extraction algorithm<br>\n✅ <strong>New capabilities</strong> - Add LLM extraction or custom strategies when needed<br>\n✅ <strong>Strategy pattern</strong> - Clean, extensible architecture</p>\n<h2 id=\"quick-start\">Quick Start</h2>\n<h3 id=\"the-simplest-way-works-like-before\">The Simplest Way (Works Like Before)</h3>\n<p>If you're already using Crawl4AI, nothing changes:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_tables</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># This works exactly like before - uses DefaultTableExtraction internally</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/data\"</span>)\n\n        <span class=\"hljs-comment\"># Tables are automatically extracted and available in result.tables</span>\n        <span class=\"hljs-keyword\">for</span> table <span class=\"hljs-keyword\">in</span> result.tables:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Table with <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(table[<span class=\"hljs-string\">'rows'</span>])}</span> rows and <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(table[<span class=\"hljs-string\">'headers'</span>])}</span> columns\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Headers: <span class=\"hljs-subst\">{table[<span class=\"hljs-string\">'headers'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"First row: <span class=\"hljs-subst\">{table[<span class=\"hljs-string\">'rows'</span>][<span class=\"hljs-number\">0</span>] <span class=\"hljs-keyword\">if</span> table[<span class=\"hljs-string\">'rows'</span>] <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">'No data'</span>}</span>\"</span>)\n\nasyncio.run(extract_tables())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"using-the-old-configuration-still-supported\">Using the Old Configuration (Still Supported)</h3>\n<p>Your existing code with <code>table_score_threshold</code> continues to work:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># This old approach STILL WORKS - we maintain backward compatibility</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=7  <span class=\"hljs-comment\"># Internally creates DefaultTableExtraction(table_score_threshold=7)</span>\n)\nresult = await crawler.arun(url, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"table-extraction-strategies_1\">Table Extraction Strategies</h2>\n<h3 id=\"understanding-the-strategy-pattern\">Understanding the Strategy Pattern</h3>\n<p>The strategy pattern allows you to choose different table extraction algorithms at runtime. Think of it as having different tools in a toolbox - you pick the right one for the job:</p>\n<ul>\n<li><strong>No explicit strategy?</strong> → Uses <code>DefaultTableExtraction</code> automatically (same as v0.7.2 and earlier)</li>\n<li><strong>Need complex table handling?</strong> → Choose <code>LLMTableExtraction</code> (costs money, use sparingly)</li>\n<li><strong>Want to disable tables?</strong> → Use <code>NoTableExtraction</code></li>\n<li><strong>Have special requirements?</strong> → Create a custom strategy</li>\n</ul>\n<h3 id=\"available-strategies\">Available Strategies</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Description</th>\n<th>Use Case</th>\n<th>Cost</th>\n<th>When to Use</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>DefaultTableExtraction</code></td>\n<td><strong>RECOMMENDED</strong>: Same algorithm as before v0.7.3</td>\n<td>General purpose (default)</td>\n<td>Free</td>\n<td><strong>Use this first - handles 95% of cases</strong></td>\n</tr>\n<tr>\n<td><code>LLMTableExtraction</code></td>\n<td>AI-powered extraction for complex tables</td>\n<td>Tables with complex rowspan/colspan</td>\n<td><strong>$$$ Per API call</strong></td>\n<td>Only when DefaultTableExtraction fails</td>\n</tr>\n<tr>\n<td><code>NoTableExtraction</code></td>\n<td>Disables table extraction</td>\n<td>When tables aren't needed</td>\n<td>Free</td>\n<td>For text-only extraction</td>\n</tr>\n<tr>\n<td>Custom strategies</td>\n<td>User-defined extraction logic</td>\n<td>Specialized requirements</td>\n<td>Free</td>\n<td>Domain-specific needs</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n<p><strong>⚠️ CRITICAL COST WARNING for LLMTableExtraction</strong>: </p>\n<p><strong>DO NOT USE <code>LLMTableExtraction</code> UNLESS ABSOLUTELY NECESSARY!</strong></p>\n<ul>\n<li><strong>Always try <code>DefaultTableExtraction</code> first</strong> - It's free and handles most tables perfectly</li>\n<li>LLM extraction <strong>costs money</strong> with every API call</li>\n<li>For large tables (100+ rows), LLM extraction can be <strong>very slow</strong></li>\n<li><strong>For large tables</strong>: If you must use LLM, choose fast providers:</li>\n<li>✅ <strong>Groq</strong> (fastest inference)</li>\n<li>✅ <strong>Cerebras</strong> (optimized for speed)</li>\n<li>⚠️ Avoid: OpenAI, Anthropic for large tables (slower)</li>\n</ul>\n<p><strong>🚧 WORK IN PROGRESS</strong>: \nWe are actively developing an <strong>advanced non-LLM algorithm</strong> that will handle complex table structures (rowspan, colspan, nested tables) for <strong>FREE</strong>. This will replace the need for costly LLM extraction in most cases. Coming soon!</p>\n</blockquote>\n<h3 id=\"defaulttableextraction\">DefaultTableExtraction</h3>\n<p>The default strategy uses a sophisticated scoring system to identify data tables:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> DefaultTableExtraction, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Customize the default extraction</span>\ntable_strategy = DefaultTableExtraction(\n    table_score_threshold=<span class=\"hljs-number\">7</span>,  <span class=\"hljs-comment\"># Scoring threshold (default: 7)</span>\n    min_rows=<span class=\"hljs-number\">2</span>,               <span class=\"hljs-comment\"># Minimum rows required</span>\n    min_cols=<span class=\"hljs-number\">2</span>,               <span class=\"hljs-comment\"># Minimum columns required</span>\n    verbose=<span class=\"hljs-literal\">True</span>              <span class=\"hljs-comment\"># Enable detailed logging</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=table_strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"scoring-system\">Scoring System</h4>\n<p>The scoring system evaluates multiple factors:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Factor</th>\n<th>Score Impact</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Has <code>&lt;thead&gt;</code></td>\n<td>+2</td>\n<td>Semantic table structure</td>\n</tr>\n<tr>\n<td>Has <code>&lt;tbody&gt;</code></td>\n<td>+1</td>\n<td>Organized table body</td>\n</tr>\n<tr>\n<td>Has <code>&lt;th&gt;</code> elements</td>\n<td>+2</td>\n<td>Header cells present</td>\n</tr>\n<tr>\n<td>Headers in correct position</td>\n<td>+1</td>\n<td>Proper semantic structure</td>\n</tr>\n<tr>\n<td>Consistent column count</td>\n<td>+2</td>\n<td>Regular data structure</td>\n</tr>\n<tr>\n<td>Has caption</td>\n<td>+2</td>\n<td>Descriptive caption</td>\n</tr>\n<tr>\n<td>Has summary</td>\n<td>+1</td>\n<td>Summary attribute</td>\n</tr>\n<tr>\n<td>High text density</td>\n<td>+2 to +3</td>\n<td>Content-rich cells</td>\n</tr>\n<tr>\n<td>Data attributes</td>\n<td>+0.5 each</td>\n<td>Data-* attributes</td>\n</tr>\n<tr>\n<td>Nested tables</td>\n<td>-3</td>\n<td>Often indicates layout</td>\n</tr>\n<tr>\n<td>Role=\"presentation\"</td>\n<td>-3</td>\n<td>Explicitly non-data</td>\n</tr>\n<tr>\n<td>Too few rows</td>\n<td>-2</td>\n<td>Insufficient data</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"llmtableextraction-use-sparingly\">LLMTableExtraction (Use Sparingly!)</h3>\n<p><strong>⚠️ WARNING</strong>: Only use this when <code>DefaultTableExtraction</code> fails with complex tables!</p>\n<p>LLMTableExtraction uses AI to understand complex table structures that traditional parsers struggle with. It automatically handles large tables through intelligent chunking and parallel processing:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMTableExtraction, LLMConfig, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Configure LLM (costs money per call!)</span>\nllm_config = LLMConfig(\n    provider=<span class=\"hljs-string\">\"groq/llama-3.3-70b-versatile\"</span>,  <span class=\"hljs-comment\"># Fast provider for large tables</span>\n    api_token=<span class=\"hljs-string\">\"your_api_key\"</span>,\n    temperature=<span class=\"hljs-number\">0.1</span>\n)\n\n<span class=\"hljs-comment\"># Create LLM extraction strategy with smart chunking</span>\ntable_strategy = LLMTableExtraction(\n    llm_config=llm_config,\n    max_tries=<span class=\"hljs-number\">3</span>,                      <span class=\"hljs-comment\"># Retry up to 3 times if extraction fails</span>\n    css_selector=<span class=\"hljs-string\">\"table\"</span>,             <span class=\"hljs-comment\"># Optional: focus on specific tables</span>\n    enable_chunking=<span class=\"hljs-literal\">True</span>,             <span class=\"hljs-comment\"># Automatically chunk large tables (default: True)</span>\n    chunk_token_threshold=<span class=\"hljs-number\">3000</span>,       <span class=\"hljs-comment\"># Split tables larger than this (default: 3000 tokens)</span>\n    min_rows_per_chunk=<span class=\"hljs-number\">10</span>,            <span class=\"hljs-comment\"># Minimum rows per chunk (default: 10)</span>\n    max_parallel_chunks=<span class=\"hljs-number\">5</span>,            <span class=\"hljs-comment\"># Process up to 5 chunks in parallel (default: 5)</span>\n    verbose=<span class=\"hljs-literal\">True</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=table_strategy\n)\n\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"when-to-use-llmtableextraction\">When to Use LLMTableExtraction</h4>\n<p>✅ <strong>Use ONLY when</strong>:\n- Tables have complex merged cells (rowspan/colspan) that break DefaultTableExtraction\n- Nested tables that need semantic understanding\n- Tables with irregular structures\n- You've tried DefaultTableExtraction and it failed</p>\n<p>❌ <strong>Never use when</strong>:\n- DefaultTableExtraction works (99% of cases)\n- Tables are simple or well-structured\n- You're processing many pages (costs add up!)\n- Tables have 100+ rows (very slow)</p>\n<h4 id=\"how-smart-chunking-works\">How Smart Chunking Works</h4>\n<p>LLMTableExtraction automatically handles large tables through intelligent chunking:</p>\n<ol>\n<li><strong>Automatic Detection</strong>: Tables exceeding the token threshold are automatically split</li>\n<li><strong>Smart Splitting</strong>: Chunks are created at row boundaries, preserving table structure</li>\n<li><strong>Header Preservation</strong>: Each chunk includes the original headers for context</li>\n<li><strong>Parallel Processing</strong>: Multiple chunks are processed simultaneously for speed</li>\n<li><strong>Intelligent Merging</strong>: Results are merged back into a single, complete table</li>\n</ol>\n<p><strong>Chunking Parameters</strong>:\n- <code>enable_chunking</code> (default: <code>True</code>): Automatically handle large tables\n- <code>chunk_token_threshold</code> (default: <code>3000</code>): When to split tables\n- <code>min_rows_per_chunk</code> (default: <code>10</code>): Ensures meaningful chunk sizes\n- <code>max_parallel_chunks</code> (default: <code>5</code>): Concurrent processing for speed</p>\n<p>The chunking is completely transparent - you get the same output format whether the table was processed in one piece or multiple chunks.</p>\n<h4 id=\"performance-optimization-for-llmtableextraction\">Performance Optimization for LLMTableExtraction</h4>\n<p><strong>Provider Recommendations by Table Size</strong>:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Table Size</th>\n<th>Recommended Providers</th>\n<th>Why</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Small (&lt;50 rows)</td>\n<td>Any provider</td>\n<td>Fast enough</td>\n</tr>\n<tr>\n<td>Medium (50-200 rows)</td>\n<td>Groq, Cerebras</td>\n<td>Optimized inference</td>\n</tr>\n<tr>\n<td>Large (200+ rows)</td>\n<td><strong>Groq</strong> (best), Cerebras</td>\n<td>Fastest inference + automatic chunking</td>\n</tr>\n<tr>\n<td>Very Large (500+ rows)</td>\n<td>Groq with chunking</td>\n<td>Parallel processing keeps it fast</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"notableextraction\">NoTableExtraction</h3>\n<p>Disable table extraction for better performance when tables aren't needed:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> NoTableExtraction, CrawlerRunConfig\n\nconfig = CrawlerRunConfig(\n    table_extraction=NoTableExtraction()\n)\n\n<span class=\"hljs-comment\"># Tables won't be extracted, improving performance</span>\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n<span class=\"hljs-keyword\">assert</span> <span class=\"hljs-built_in\">len</span>(result.tables) == <span class=\"hljs-number\">0</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"extracted-table-structure\">Extracted Table Structure</h2>\n<p>Each extracted table contains:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"headers\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"Column 1\"</span>, <span class=\"hljs-string\">\"Column 2\"</span>, <span class=\"hljs-punctuation\">...</span><span class=\"hljs-punctuation\">]</span>,  <span class=\"hljs-comment\"># Column headers</span>\n    <span class=\"hljs-string\">\"rows\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>                                   <span class=\"hljs-comment\"># Data rows</span>\n        <span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"Row 1 Col 1\"</span>, <span class=\"hljs-string\">\"Row 1 Col 2\"</span>, <span class=\"hljs-punctuation\">...</span><span class=\"hljs-punctuation\">]</span>,\n        <span class=\"hljs-punctuation\">[</span><span class=\"hljs-string\">\"Row 2 Col 1\"</span>, <span class=\"hljs-string\">\"Row 2 Col 2\"</span>, <span class=\"hljs-punctuation\">...</span><span class=\"hljs-punctuation\">]</span>,\n    <span class=\"hljs-punctuation\">]</span>,\n    <span class=\"hljs-string\">\"caption\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Table Caption\"</span>,                <span class=\"hljs-comment\"># If present</span>\n    <span class=\"hljs-string\">\"summary\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Table Summary\"</span>,                <span class=\"hljs-comment\"># If present</span>\n    <span class=\"hljs-string\">\"metadata\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">{</span>\n        <span class=\"hljs-string\">\"row_count\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">10</span>,                       <span class=\"hljs-comment\"># Number of rows</span>\n        <span class=\"hljs-string\">\"column_count\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">3</span>,                      <span class=\"hljs-comment\"># Number of columns</span>\n        <span class=\"hljs-string\">\"has_headers\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Headers detected</span>\n        <span class=\"hljs-string\">\"has_caption\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\">True</span>,                    <span class=\"hljs-comment\"># Caption exists</span>\n        <span class=\"hljs-string\">\"has_summary\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\">False</span>,                   <span class=\"hljs-comment\"># Summary exists</span>\n        <span class=\"hljs-string\">\"id\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"data-table-1\"</span>,                   <span class=\"hljs-comment\"># Table ID if present</span>\n        <span class=\"hljs-string\">\"class\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"financial-data\"</span>               <span class=\"hljs-comment\"># Table class if present</span>\n    <span class=\"hljs-punctuation\">}</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"configuration-options\">Configuration Options</h2>\n<h3 id=\"basic-configuration\">Basic Configuration</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">config = CrawlerRunConfig(\n    <span class=\"hljs-comment\"># Table extraction settings</span>\n    table_score_threshold=7,      <span class=\"hljs-comment\"># Default threshold (backward compatible)</span>\n    table_extraction=strategy,     <span class=\"hljs-comment\"># Optional: custom strategy</span>\n\n    <span class=\"hljs-comment\"># Filter what to process</span>\n    css_selector=<span class=\"hljs-string\">\"main\"</span>,          <span class=\"hljs-comment\"># Focus on specific area</span>\n    excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"aside\"</span>] <span class=\"hljs-comment\"># Exclude page sections</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"advanced-configuration\">Advanced Configuration</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> DefaultTableExtraction, CrawlerRunConfig\n\n<span class=\"hljs-comment\"># Fine-tuned extraction</span>\nstrategy = DefaultTableExtraction(\n    table_score_threshold=<span class=\"hljs-number\">5</span>,      <span class=\"hljs-comment\"># Lower = more permissive</span>\n    min_rows=<span class=\"hljs-number\">3</span>,                   <span class=\"hljs-comment\"># Require at least 3 rows</span>\n    min_cols=<span class=\"hljs-number\">2</span>,                   <span class=\"hljs-comment\"># Require at least 2 columns</span>\n    verbose=<span class=\"hljs-literal\">True</span>                  <span class=\"hljs-comment\"># Detailed logging</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=strategy,\n    css_selector=<span class=\"hljs-string\">\"article.content\"</span>, <span class=\"hljs-comment\"># Target specific content</span>\n    exclude_domains=[<span class=\"hljs-string\">\"ads.com\"</span>],   <span class=\"hljs-comment\"># Exclude ad domains</span>\n    cache_mode=CacheMode.BYPASS    <span class=\"hljs-comment\"># Fresh extraction</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"working-with-extracted-tables\">Working with Extracted Tables</h2>\n<h3 id=\"convert-to-pandas-dataframe\">Convert to Pandas DataFrame</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-function\">import pandas <span class=\"hljs-keyword\">as</span> pd\n\n<span class=\"hljs-keyword\">async</span> def <span class=\"hljs-title\">tables_to_dataframes</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n        result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(url)\n\n        dataframes = []\n        <span class=\"hljs-keyword\">for</span> table_data <span class=\"hljs-keyword\">in</span> result.tables:\n            <span class=\"hljs-meta\"># Create DataFrame</span>\n            <span class=\"hljs-keyword\">if</span> table_data[<span class=\"hljs-string\">'headers'</span>]:\n                df = pd.DataFrame(\n                    table_data[<span class=\"hljs-string\">'rows'</span>],\n                    columns=table_data[<span class=\"hljs-string\">'headers'</span>]\n                )\n            <span class=\"hljs-keyword\">else</span>:\n                df = pd.DataFrame(table_data[<span class=\"hljs-string\">'rows'</span>])\n\n            <span class=\"hljs-meta\"># Add metadata as DataFrame attributes</span>\n            df.attrs[<span class=\"hljs-string\">'caption'</span>] = table_data.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">'caption'</span>, <span class=\"hljs-string\">''</span>)\n            df.attrs[<span class=\"hljs-string\">'metadata'</span>] = table_data.<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">'metadata'</span>, {})\n\n            dataframes.append(df)\n\n        <span class=\"hljs-keyword\">return</span> dataframes\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"filter-tables-by-criteria\">Filter Tables by Criteria</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_large_tables</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Configure minimum size requirements</span>\n        strategy = DefaultTableExtraction(\n            min_rows=<span class=\"hljs-number\">10</span>,\n            min_cols=<span class=\"hljs-number\">3</span>,\n            table_score_threshold=<span class=\"hljs-number\">6</span>\n        )\n\n        config = CrawlerRunConfig(\n            table_extraction=strategy\n        )\n\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n\n        <span class=\"hljs-comment\"># Further filter results</span>\n        large_tables = [\n            table <span class=\"hljs-keyword\">for</span> table <span class=\"hljs-keyword\">in</span> result.tables\n            <span class=\"hljs-keyword\">if</span> table[<span class=\"hljs-string\">'metadata'</span>][<span class=\"hljs-string\">'row_count'</span>] &gt; <span class=\"hljs-number\">10</span>\n            <span class=\"hljs-keyword\">and</span> table[<span class=\"hljs-string\">'metadata'</span>][<span class=\"hljs-string\">'column_count'</span>] &gt; <span class=\"hljs-number\">3</span>\n        ]\n\n        <span class=\"hljs-keyword\">return</span> large_tables\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"export-tables-to-different-formats\">Export Tables to Different Formats</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> csv\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">export_tables</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url)\n\n        <span class=\"hljs-keyword\">for</span> i, table <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(result.tables):\n            <span class=\"hljs-comment\"># Export as JSON</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">f'table_<span class=\"hljs-subst\">{i}</span>.json'</span>, <span class=\"hljs-string\">'w'</span>) <span class=\"hljs-keyword\">as</span> f:\n                json.dump(table, f, indent=<span class=\"hljs-number\">2</span>)\n\n            <span class=\"hljs-comment\"># Export as CSV</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">f'table_<span class=\"hljs-subst\">{i}</span>.csv'</span>, <span class=\"hljs-string\">'w'</span>, newline=<span class=\"hljs-string\">''</span>) <span class=\"hljs-keyword\">as</span> f:\n                writer = csv.writer(f)\n                <span class=\"hljs-keyword\">if</span> table[<span class=\"hljs-string\">'headers'</span>]:\n                    writer.writerow(table[<span class=\"hljs-string\">'headers'</span>])\n                writer.writerows(table[<span class=\"hljs-string\">'rows'</span>])\n\n            <span class=\"hljs-comment\"># Export as Markdown</span>\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">f'table_<span class=\"hljs-subst\">{i}</span>.md'</span>, <span class=\"hljs-string\">'w'</span>) <span class=\"hljs-keyword\">as</span> f:\n                <span class=\"hljs-comment\"># Write headers</span>\n                <span class=\"hljs-keyword\">if</span> table[<span class=\"hljs-string\">'headers'</span>]:\n                    f.write(<span class=\"hljs-string\">'| '</span> + <span class=\"hljs-string\">' | '</span>.join(table[<span class=\"hljs-string\">'headers'</span>]) + <span class=\"hljs-string\">' |\\n'</span>)\n                    f.write(<span class=\"hljs-string\">'|'</span> + <span class=\"hljs-string\">'---|'</span> * <span class=\"hljs-built_in\">len</span>(table[<span class=\"hljs-string\">'headers'</span>]) + <span class=\"hljs-string\">'\\n'</span>)\n\n                <span class=\"hljs-comment\"># Write rows</span>\n                <span class=\"hljs-keyword\">for</span> row <span class=\"hljs-keyword\">in</span> table[<span class=\"hljs-string\">'rows'</span>]:\n                    f.write(<span class=\"hljs-string\">'| '</span> + <span class=\"hljs-string\">' | '</span>.join(<span class=\"hljs-built_in\">str</span>(cell) <span class=\"hljs-keyword\">for</span> cell <span class=\"hljs-keyword\">in</span> row) + <span class=\"hljs-string\">' |\\n'</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"creating-custom-strategies\">Creating Custom Strategies</h2>\n<p>Extend <code>TableExtractionStrategy</code> to create custom extraction logic:</p>\n<h3 id=\"example-financial-table-extractor\">Example: Financial Table Extractor</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> TableExtractionStrategy\n<span class=\"hljs-keyword\">from</span> typing <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">List</span>, <span class=\"hljs-type\">Dict</span>, <span class=\"hljs-type\">Any</span>\n<span class=\"hljs-keyword\">import</span> re\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">FinancialTableExtractor</span>(<span class=\"hljs-title class_ inherited__\">TableExtractionStrategy</span>):\n    <span class=\"hljs-string\">\"\"\"Extract tables containing financial data.\"\"\"</span>\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, currency_symbols=<span class=\"hljs-literal\">None</span>, require_numbers=<span class=\"hljs-literal\">True</span>, **kwargs</span>):\n        <span class=\"hljs-built_in\">super</span>().__init__(**kwargs)\n        self.currency_symbols = currency_symbols <span class=\"hljs-keyword\">or</span> [<span class=\"hljs-string\">'$'</span>, <span class=\"hljs-string\">'€'</span>, <span class=\"hljs-string\">'£'</span>, <span class=\"hljs-string\">'¥'</span>]\n        self.require_numbers = require_numbers\n        self.number_pattern = re.<span class=\"hljs-built_in\">compile</span>(<span class=\"hljs-string\">r'\\d+[,.]?\\d*'</span>)\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_tables</span>(<span class=\"hljs-params\">self, element, **kwargs</span>):\n        tables_data = []\n\n        <span class=\"hljs-keyword\">for</span> table <span class=\"hljs-keyword\">in</span> element.xpath(<span class=\"hljs-string\">\".//table\"</span>):\n            <span class=\"hljs-comment\"># Check if table contains financial indicators</span>\n            table_text = <span class=\"hljs-string\">''</span>.join(table.itertext())\n\n            <span class=\"hljs-comment\"># Must contain currency symbols</span>\n            has_currency = <span class=\"hljs-built_in\">any</span>(sym <span class=\"hljs-keyword\">in</span> table_text <span class=\"hljs-keyword\">for</span> sym <span class=\"hljs-keyword\">in</span> self.currency_symbols)\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> has_currency:\n                <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-comment\"># Must contain numbers if required</span>\n            <span class=\"hljs-keyword\">if</span> self.require_numbers:\n                numbers = self.number_pattern.findall(table_text)\n                <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">len</span>(numbers) &lt; <span class=\"hljs-number\">3</span>:  <span class=\"hljs-comment\"># Arbitrary minimum</span>\n                    <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-comment\"># Extract the table data</span>\n            table_data = self._extract_financial_data(table)\n            <span class=\"hljs-keyword\">if</span> table_data:\n                tables_data.append(table_data)\n\n        <span class=\"hljs-keyword\">return</span> tables_data\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">_extract_financial_data</span>(<span class=\"hljs-params\">self, table</span>):\n        <span class=\"hljs-string\">\"\"\"Extract and clean financial data from table.\"\"\"</span>\n        headers = []\n        rows = []\n\n        <span class=\"hljs-comment\"># Extract headers</span>\n        <span class=\"hljs-keyword\">for</span> th <span class=\"hljs-keyword\">in</span> table.xpath(<span class=\"hljs-string\">\".//thead//th | .//tr[1]//th\"</span>):\n            headers.append(th.text_content().strip())\n\n        <span class=\"hljs-comment\"># Extract and clean rows</span>\n        <span class=\"hljs-keyword\">for</span> tr <span class=\"hljs-keyword\">in</span> table.xpath(<span class=\"hljs-string\">\".//tbody//tr | .//tr[position()&gt;1]\"</span>):\n            row = []\n            <span class=\"hljs-keyword\">for</span> td <span class=\"hljs-keyword\">in</span> tr.xpath(<span class=\"hljs-string\">\".//td\"</span>):\n                text = td.text_content().strip()\n                <span class=\"hljs-comment\"># Clean currency formatting</span>\n                text = re.sub(<span class=\"hljs-string\">r'[$€£¥,]'</span>, <span class=\"hljs-string\">''</span>, text)\n                row.append(text)\n            <span class=\"hljs-keyword\">if</span> row:\n                rows.append(row)\n\n        <span class=\"hljs-keyword\">return</span> {\n            <span class=\"hljs-string\">\"headers\"</span>: headers,\n            <span class=\"hljs-string\">\"rows\"</span>: rows,\n            <span class=\"hljs-string\">\"caption\"</span>: self._get_caption(table),\n            <span class=\"hljs-string\">\"summary\"</span>: table.get(<span class=\"hljs-string\">\"summary\"</span>, <span class=\"hljs-string\">\"\"</span>),\n            <span class=\"hljs-string\">\"metadata\"</span>: {\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"financial\"</span>,\n                <span class=\"hljs-string\">\"row_count\"</span>: <span class=\"hljs-built_in\">len</span>(rows),\n                <span class=\"hljs-string\">\"column_count\"</span>: <span class=\"hljs-built_in\">len</span>(headers) <span class=\"hljs-keyword\">or</span> <span class=\"hljs-built_in\">len</span>(rows[<span class=\"hljs-number\">0</span>]) <span class=\"hljs-keyword\">if</span> rows <span class=\"hljs-keyword\">else</span> <span class=\"hljs-number\">0</span>\n            }\n        }\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">_get_caption</span>(<span class=\"hljs-params\">self, table</span>):\n        caption = table.xpath(<span class=\"hljs-string\">\".//caption/text()\"</span>)\n        <span class=\"hljs-keyword\">return</span> caption[<span class=\"hljs-number\">0</span>].strip() <span class=\"hljs-keyword\">if</span> caption <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">\"\"</span>\n\n<span class=\"hljs-comment\"># Usage</span>\nstrategy = FinancialTableExtractor(\n    currency_symbols=[<span class=\"hljs-string\">'$'</span>, <span class=\"hljs-string\">'EUR'</span>],\n    require_numbers=<span class=\"hljs-literal\">True</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"example-specific-table-extractor\">Example: Specific Table Extractor</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">SpecificTableExtractor</span>(<span class=\"hljs-title class_ inherited__\">TableExtractionStrategy</span>):\n    <span class=\"hljs-string\">\"\"\"Extract only tables matching specific criteria.\"\"\"</span>\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, \n                 required_headers=<span class=\"hljs-literal\">None</span>, \n                 id_pattern=<span class=\"hljs-literal\">None</span>,\n                 class_pattern=<span class=\"hljs-literal\">None</span>,\n                 **kwargs</span>):\n        <span class=\"hljs-built_in\">super</span>().__init__(**kwargs)\n        self.required_headers = required_headers <span class=\"hljs-keyword\">or</span> []\n        self.id_pattern = id_pattern\n        self.class_pattern = class_pattern\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_tables</span>(<span class=\"hljs-params\">self, element, **kwargs</span>):\n        tables_data = []\n\n        <span class=\"hljs-keyword\">for</span> table <span class=\"hljs-keyword\">in</span> element.xpath(<span class=\"hljs-string\">\".//table\"</span>):\n            <span class=\"hljs-comment\"># Check ID pattern</span>\n            <span class=\"hljs-keyword\">if</span> self.id_pattern:\n                table_id = table.get(<span class=\"hljs-string\">'id'</span>, <span class=\"hljs-string\">''</span>)\n                <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> re.<span class=\"hljs-keyword\">match</span>(self.id_pattern, table_id):\n                    <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-comment\"># Check class pattern</span>\n            <span class=\"hljs-keyword\">if</span> self.class_pattern:\n                table_class = table.get(<span class=\"hljs-string\">'class'</span>, <span class=\"hljs-string\">''</span>)\n                <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> re.<span class=\"hljs-keyword\">match</span>(self.class_pattern, table_class):\n                    <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-comment\"># Extract headers to check requirements</span>\n            headers = self._extract_headers(table)\n\n            <span class=\"hljs-comment\"># Check if required headers are present</span>\n            <span class=\"hljs-keyword\">if</span> self.required_headers:\n                <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> <span class=\"hljs-built_in\">all</span>(req <span class=\"hljs-keyword\">in</span> headers <span class=\"hljs-keyword\">for</span> req <span class=\"hljs-keyword\">in</span> self.required_headers):\n                    <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-comment\"># Extract full table data</span>\n            table_data = self._extract_table_data(table, headers)\n            tables_data.append(table_data)\n\n        <span class=\"hljs-keyword\">return</span> tables_data\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"combining-with-other-strategies\">Combining with Other Strategies</h2>\n<p>Table extraction works seamlessly with other Crawl4AI strategies:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    DefaultTableExtraction,\n    LLMExtractionStrategy,\n    JsonCssExtractionStrategy\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">combined_extraction</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        config = CrawlerRunConfig(\n            <span class=\"hljs-comment\"># Table extraction</span>\n            table_extraction=DefaultTableExtraction(\n                table_score_threshold=<span class=\"hljs-number\">6</span>,\n                min_rows=<span class=\"hljs-number\">2</span>\n            ),\n\n            <span class=\"hljs-comment\"># CSS-based extraction for specific elements</span>\n            extraction_strategy=JsonCssExtractionStrategy({\n                <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"h1\"</span>,\n                <span class=\"hljs-string\">\"summary\"</span>: <span class=\"hljs-string\">\"p.summary\"</span>,\n                <span class=\"hljs-string\">\"date\"</span>: <span class=\"hljs-string\">\"time\"</span>\n            }),\n\n            <span class=\"hljs-comment\"># Focus on main content</span>\n            css_selector=<span class=\"hljs-string\">\"main.content\"</span>\n        )\n\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n\n        <span class=\"hljs-comment\"># Access different extraction results</span>\n        tables = result.tables  <span class=\"hljs-comment\"># Table data</span>\n        structured = json.loads(result.extracted_content)  <span class=\"hljs-comment\"># CSS extraction</span>\n\n        <span class=\"hljs-keyword\">return</span> {\n            <span class=\"hljs-string\">\"tables\"</span>: tables,\n            <span class=\"hljs-string\">\"structured_data\"</span>: structured,\n            <span class=\"hljs-string\">\"markdown\"</span>: result.markdown\n        }\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"performance-considerations\">Performance Considerations</h2>\n<h3 id=\"optimization-tips\">Optimization Tips</h3>\n<ol>\n<li><strong>Disable when not needed</strong>: Use <code>NoTableExtraction</code> if tables aren't required</li>\n<li><strong>Target specific areas</strong>: Use <code>css_selector</code> to limit processing scope</li>\n<li><strong>Set minimum thresholds</strong>: Filter out small/irrelevant tables early</li>\n<li><strong>Cache results</strong>: Use appropriate cache modes for repeated extractions</li>\n</ol>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Optimized configuration for large pages</span>\nconfig = CrawlerRunConfig(\n    <span class=\"hljs-comment\"># Only process main content area</span>\n    css_selector=<span class=\"hljs-string\">\"article.main-content\"</span>,\n\n    <span class=\"hljs-comment\"># Exclude navigation and sidebars</span>\n    excluded_tags=[<span class=\"hljs-string\">\"nav\"</span>, <span class=\"hljs-string\">\"aside\"</span>, <span class=\"hljs-string\">\"footer\"</span>],\n\n    <span class=\"hljs-comment\"># Higher threshold for stricter filtering</span>\n    table_extraction=DefaultTableExtraction(\n        table_score_threshold=8,\n        min_rows=5,\n        min_cols=3\n    ),\n\n    <span class=\"hljs-comment\"># Enable caching for repeated access</span>\n    cache_mode=CacheMode.ENABLED\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"migration-guide\">Migration Guide</h2>\n<h3 id=\"important-your-code-still-works\">Important: Your Code Still Works!</h3>\n<p><strong>No changes required!</strong> The transition to the strategy pattern is <strong>fully backward compatible</strong>.</p>\n<h3 id=\"how-it-works-internally\">How It Works Internally</h3>\n<h4 id=\"v072-and-earlier\">v0.7.2 and Earlier</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Old way - directly passing table_score_threshold</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=7\n)\n<span class=\"hljs-comment\"># Internally: No strategy pattern, direct implementation</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"v073-current\">v0.7.3+ (Current)</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Old way STILL WORKS - we handle it internally</span>\nconfig <span class=\"hljs-punctuation\">=</span> CrawlerRunConfig<span class=\"hljs-punctuation\">(</span>\n    table_score_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">7</span>\n<span class=\"hljs-punctuation\">)</span>\n<span class=\"hljs-comment\"># Internally: Automatically creates DefaultTableExtraction(table_score_threshold=7)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"taking-advantage-of-new-features\">Taking Advantage of New Features</h3>\n<p>While your old code works, you can now use the strategy pattern for more control:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Option 1: Keep using the old way (perfectly fine!)</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=7  <span class=\"hljs-comment\"># Still supported</span>\n)\n\n<span class=\"hljs-comment\"># Option 2: Use the new strategy pattern (more flexibility)</span>\nfrom crawl4ai import DefaultTableExtraction\n\nstrategy = DefaultTableExtraction(\n    table_score_threshold=7,\n    min_rows=2,  <span class=\"hljs-comment\"># New capability!</span>\n    min_cols=2   <span class=\"hljs-comment\"># New capability!</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=strategy\n)\n\n<span class=\"hljs-comment\"># Option 3: Use advanced strategies when needed</span>\nfrom crawl4ai import LLMTableExtraction, LLMConfig\n\n<span class=\"hljs-comment\"># Only for complex tables that DefaultTableExtraction can't handle</span>\n<span class=\"hljs-comment\"># Automatically handles large tables with smart chunking</span>\nllm_strategy = LLMTableExtraction(\n    llm_config=LLMConfig(\n        provider=<span class=\"hljs-string\">\"groq/llama-3.3-70b-versatile\"</span>,\n        api_token=<span class=\"hljs-string\">\"your_key\"</span>\n    ),\n    max_tries=3,\n    enable_chunking=True,  <span class=\"hljs-comment\"># Automatically chunk large tables</span>\n    chunk_token_threshold=3000,  <span class=\"hljs-comment\"># Chunk when exceeding 3000 tokens</span>\n    max_parallel_chunks=5  <span class=\"hljs-comment\"># Process up to 5 chunks in parallel</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=llm_strategy  <span class=\"hljs-comment\"># Advanced extraction with automatic chunking</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"summary\">Summary</h3>\n<ul>\n<li>✅ <strong>No breaking changes</strong> - Old code works as-is</li>\n<li>✅ <strong>Same defaults</strong> - DefaultTableExtraction is automatically used</li>\n<li>✅ <strong>Gradual adoption</strong> - Use new features when you need them</li>\n<li>✅ <strong>Full compatibility</strong> - result.tables structure unchanged</li>\n</ul>\n<h2 id=\"best-practices\">Best Practices</h2>\n<h3 id=\"1-choose-the-right-strategy-cost-conscious-approach\">1. Choose the Right Strategy (Cost-Conscious Approach)</h3>\n<p><strong>Decision Flow</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"><span class=\"hljs-number\">1.</span> Do you need tables? \n   → <span class=\"hljs-keyword\">No</span>: Use NoTableExtraction\n   → Yes: Continue <span class=\"hljs-keyword\">to</span> #<span class=\"hljs-number\">2</span>\n\n<span class=\"hljs-number\">2.</span> Try DefaultTableExtraction <span class=\"hljs-keyword\">first</span> (<span class=\"hljs-keyword\">FREE</span>)\n   → Works? Done<span class=\"hljs-operator\">!</span> ✅\n   → Fails? Continue <span class=\"hljs-keyword\">to</span> #<span class=\"hljs-number\">3</span>\n\n<span class=\"hljs-number\">3.</span> <span class=\"hljs-keyword\">Is</span> the <span class=\"hljs-keyword\">table</span> critical <span class=\"hljs-keyword\">and</span> complex?\n   → <span class=\"hljs-keyword\">No</span>: Accept DefaultTableExtraction results\n   → Yes: Continue <span class=\"hljs-keyword\">to</span> #<span class=\"hljs-number\">4</span>\n\n<span class=\"hljs-number\">4.</span> Use LLMTableExtraction (COSTS MONEY)\n   → Small <span class=\"hljs-keyword\">table</span> (<span class=\"hljs-operator\">&lt;</span><span class=\"hljs-number\">50</span> <span class=\"hljs-keyword\">rows</span>): <span class=\"hljs-keyword\">Any</span> LLM provider\n   → <span class=\"hljs-keyword\">Large</span> <span class=\"hljs-keyword\">table</span> (<span class=\"hljs-number\">50</span><span class=\"hljs-operator\">+</span> <span class=\"hljs-keyword\">rows</span>): Use Groq <span class=\"hljs-keyword\">or</span> Cerebras\n   → Very <span class=\"hljs-keyword\">large</span> (<span class=\"hljs-number\">500</span><span class=\"hljs-operator\">+</span> <span class=\"hljs-keyword\">rows</span>): Reconsider <span class=\"hljs-operator\">-</span> maybe chunk the page\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>Strategy Selection Guide</strong>:\n- <strong>DefaultTableExtraction</strong>: Use for 99% of cases - it's free and effective\n- <strong>LLMTableExtraction</strong>: Only for complex tables with merged cells that break DefaultTableExtraction\n- <strong>NoTableExtraction</strong>: When you only need text/markdown content\n- <strong>Custom Strategy</strong>: For specialized requirements (financial, scientific, etc.)</p>\n<h3 id=\"2-validate-extracted-data\">2. Validate Extracted Data</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">validate_table</span>(<span class=\"hljs-params\">table</span>):\n    <span class=\"hljs-string\">\"\"\"Validate table data quality.\"\"\"</span>\n    <span class=\"hljs-comment\"># Check structure</span>\n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> table.get(<span class=\"hljs-string\">'rows'</span>):\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">False</span>\n\n    <span class=\"hljs-comment\"># Check consistency</span>\n    <span class=\"hljs-keyword\">if</span> table.get(<span class=\"hljs-string\">'headers'</span>):\n        expected_cols = <span class=\"hljs-built_in\">len</span>(table[<span class=\"hljs-string\">'headers'</span>])\n        <span class=\"hljs-keyword\">for</span> row <span class=\"hljs-keyword\">in</span> table[<span class=\"hljs-string\">'rows'</span>]:\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">len</span>(row) != expected_cols:\n                <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">False</span>\n\n    <span class=\"hljs-comment\"># Check minimum content</span>\n    total_cells = <span class=\"hljs-built_in\">sum</span>(<span class=\"hljs-built_in\">len</span>(row) <span class=\"hljs-keyword\">for</span> row <span class=\"hljs-keyword\">in</span> table[<span class=\"hljs-string\">'rows'</span>])\n    non_empty = <span class=\"hljs-built_in\">sum</span>(<span class=\"hljs-number\">1</span> <span class=\"hljs-keyword\">for</span> row <span class=\"hljs-keyword\">in</span> table[<span class=\"hljs-string\">'rows'</span>] \n                    <span class=\"hljs-keyword\">for</span> cell <span class=\"hljs-keyword\">in</span> row <span class=\"hljs-keyword\">if</span> cell.strip())\n\n    <span class=\"hljs-keyword\">if</span> non_empty / total_cells &lt; <span class=\"hljs-number\">0.5</span>:  <span class=\"hljs-comment\"># Less than 50% non-empty</span>\n        <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">False</span>\n\n    <span class=\"hljs-keyword\">return</span> <span class=\"hljs-literal\">True</span>\n\n<span class=\"hljs-comment\"># Filter valid tables</span>\nvalid_tables = [t <span class=\"hljs-keyword\">for</span> t <span class=\"hljs-keyword\">in</span> result.tables <span class=\"hljs-keyword\">if</span> validate_table(t)]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-handle-edge-cases\">3. Handle Edge Cases</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">robust_table_extraction</span>(<span class=\"hljs-params\">url</span>):\n    <span class=\"hljs-string\">\"\"\"Extract tables with error handling.\"\"\"</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-keyword\">try</span>:\n            config = CrawlerRunConfig(\n                table_extraction=DefaultTableExtraction(\n                    table_score_threshold=<span class=\"hljs-number\">6</span>,\n                    verbose=<span class=\"hljs-literal\">True</span>\n                )\n            )\n\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Crawl failed: <span class=\"hljs-subst\">{result.error}</span>\"</span>)\n                <span class=\"hljs-keyword\">return</span> []\n\n            <span class=\"hljs-comment\"># Process tables safely</span>\n            processed_tables = []\n            <span class=\"hljs-keyword\">for</span> table <span class=\"hljs-keyword\">in</span> result.tables:\n                <span class=\"hljs-keyword\">try</span>:\n                    <span class=\"hljs-comment\"># Validate and process</span>\n                    <span class=\"hljs-keyword\">if</span> validate_table(table):\n                        processed_tables.append(table)\n                <span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error processing table: <span class=\"hljs-subst\">{e}</span>\"</span>)\n                    <span class=\"hljs-keyword\">continue</span>\n\n            <span class=\"hljs-keyword\">return</span> processed_tables\n\n        <span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extraction error: <span class=\"hljs-subst\">{e}</span>\"</span>)\n            <span class=\"hljs-keyword\">return</span> []\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"troubleshooting\">Troubleshooting</h2>\n<h3 id=\"common-issues-and-solutions\">Common Issues and Solutions</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Issue</th>\n<th>Cause</th>\n<th>Solution</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No tables extracted</td>\n<td>Score too high</td>\n<td>Lower <code>table_score_threshold</code></td>\n</tr>\n<tr>\n<td>Layout tables included</td>\n<td>Score too low</td>\n<td>Increase <code>table_score_threshold</code></td>\n</tr>\n<tr>\n<td>Missing tables</td>\n<td>CSS selector too specific</td>\n<td>Broaden or remove <code>css_selector</code></td>\n</tr>\n<tr>\n<td>Incomplete data</td>\n<td>Complex table structure</td>\n<td>Create custom strategy</td>\n</tr>\n<tr>\n<td>Performance issues</td>\n<td>Processing entire page</td>\n<td>Use <code>css_selector</code> to limit scope</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"debug-logging\">Debug Logging</h3>\n<p>Enable verbose logging to understand extraction decisions:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> logging\n\n<span class=\"hljs-comment\"># Configure logging</span>\nlogging.basicConfig(level=logging.DEBUG)\n\n<span class=\"hljs-comment\"># Enable verbose mode in strategy</span>\nstrategy = DefaultTableExtraction(\n    table_score_threshold=<span class=\"hljs-number\">7</span>,\n    verbose=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Detailed extraction logs</span>\n)\n\nconfig = CrawlerRunConfig(\n    table_extraction=strategy,\n    verbose=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># General crawler logs</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"see-also\">See Also</h2>\n<ul>\n<li><a href=\"extraction-strategies.md\">Extraction Strategies</a> - Overview of all extraction strategies</li>\n<li><a href=\"../content-selection/\">Content Selection</a> - Using CSS selectors and filters</li>\n<li><a href=\"../optimization/performance-tuning.md\">Performance Optimization</a> - Speed up extraction</li>\n<li><a href=\"../examples/table_extraction_example.py\">Examples</a> - Complete working examples</li>\n</ul>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/core/url-seeding/",
    "title": "URL Seeding - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"url-seeding-the-smart-way-to-crawl-at-scale\">URL Seeding: The Smart Way to Crawl at Scale</h1>\n<h2 id=\"why-url-seeding\">Why URL Seeding?</h2>\n<p>Web crawling comes in different flavors, each with its own strengths. Let's understand when to use URL seeding versus deep crawling.</p>\n<h3 id=\"deep-crawling-real-time-discovery\">Deep Crawling: Real-Time Discovery</h3>\n<p>Deep crawling is perfect when you need:\n- <strong>Fresh, real-time data</strong> - discovering pages as they're created\n- <strong>Dynamic exploration</strong> - following links based on content\n- <strong>Selective extraction</strong> - stopping when you find what you need</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Deep crawling example: Explore a website dynamically</span>\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai.deep_crawling <span class=\"hljs-keyword\">import</span> BFSDeepCrawlStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">deep_crawl_example</span>():\n    <span class=\"hljs-comment\"># Configure a 2-level deep crawl</span>\n    config = CrawlerRunConfig(\n        deep_crawl_strategy=BFSDeepCrawlStrategy(\n            max_depth=<span class=\"hljs-number\">2</span>,           <span class=\"hljs-comment\"># Crawl 2 levels deep</span>\n            include_external=<span class=\"hljs-literal\">False</span>, <span class=\"hljs-comment\"># Stay within domain</span>\n            max_pages=<span class=\"hljs-number\">50</span>           <span class=\"hljs-comment\"># Limit for efficiency</span>\n        ),\n        verbose=<span class=\"hljs-literal\">True</span>\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Start crawling and follow links dynamically</span>\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com\"</span>, config=config)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Discovered and crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> pages\"</span>)\n        <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results[:<span class=\"hljs-number\">3</span>]:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found: <span class=\"hljs-subst\">{result.url}</span> at depth <span class=\"hljs-subst\">{result.metadata.get(<span class=\"hljs-string\">'depth'</span>, <span class=\"hljs-number\">0</span>)}</span>\"</span>)\n\nasyncio.run(deep_crawl_example())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"url-seeding-bulk-discovery\">URL Seeding: Bulk Discovery</h3>\n<p>URL seeding shines when you want:\n- <strong>Comprehensive coverage</strong> - get thousands of URLs in seconds\n- <strong>Bulk processing</strong> - filter before crawling\n- <strong>Resource efficiency</strong> - know exactly what you'll crawl</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># URL seeding example: Analyze all documentation</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncUrlSeeder, SeedingConfig\n\nseeder = AsyncUrlSeeder()\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    pattern=<span class=\"hljs-string\">\"*/docs/*\"</span>\n)\n\n<span class=\"hljs-comment\"># Get ALL documentation URLs instantly</span>\nurls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n<span class=\"hljs-comment\"># 1000+ URLs discovered in seconds!</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"the-trade-offs\">The Trade-offs</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Aspect</th>\n<th>Deep Crawling</th>\n<th>URL Seeding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Coverage</strong></td>\n<td>Discovers pages dynamically</td>\n<td>Gets most existing URLs instantly</td>\n</tr>\n<tr>\n<td><strong>Freshness</strong></td>\n<td>Finds brand new pages</td>\n<td>May miss very recent pages</td>\n</tr>\n<tr>\n<td><strong>Speed</strong></td>\n<td>Slower, page by page</td>\n<td>Extremely fast bulk discovery</td>\n</tr>\n<tr>\n<td><strong>Resource Usage</strong></td>\n<td>Higher - crawls to discover</td>\n<td>Lower - discovers then crawls</td>\n</tr>\n<tr>\n<td><strong>Control</strong></td>\n<td>Can stop mid-process</td>\n<td>Pre-filters before crawling</td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"when-to-use-each\">When to Use Each</h3>\n<p><strong>Choose Deep Crawling when:</strong>\n- You need the absolute latest content\n- You're searching for specific information\n- The site structure is unknown or dynamic\n- You want to stop as soon as you find what you need</p>\n<p><strong>Choose URL Seeding when:</strong>\n- You need to analyze large portions of a site\n- You want to filter URLs before crawling\n- You're doing comparative analysis\n- You need to optimize resource usage</p>\n<p>The magic happens when you understand both approaches and choose the right tool for your task. Sometimes, you might even combine them - use URL seeding for bulk discovery, then deep crawl specific sections for the latest updates.</p>\n<h2 id=\"your-first-url-seeding-adventure\">Your First URL Seeding Adventure</h2>\n<p>Let's see the magic in action. We'll discover blog posts about Python, filter for tutorials, and crawl only those pages.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncUrlSeeder, AsyncWebCrawler, SeedingConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">smart_blog_crawler</span>():\n    <span class=\"hljs-comment\"># Step 1: Create our URL discoverer</span>\n    seeder = AsyncUrlSeeder()\n\n    <span class=\"hljs-comment\"># Step 2: Configure discovery - let's find all blog posts</span>\n    config = SeedingConfig(\n        source=<span class=\"hljs-string\">\"sitemap+cc\"</span>,      <span class=\"hljs-comment\"># Use the website's sitemap+cc</span>\n        pattern=<span class=\"hljs-string\">\"*/courses/*\"</span>,    <span class=\"hljs-comment\"># Only courses related posts</span>\n        extract_head=<span class=\"hljs-literal\">True</span>,          <span class=\"hljs-comment\"># Get page metadata</span>\n        max_urls=<span class=\"hljs-number\">100</span>               <span class=\"hljs-comment\"># Limit for this example</span>\n    )\n\n    <span class=\"hljs-comment\"># Step 3: Discover URLs from the Python blog</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"🔍 Discovering course posts...\"</span>)\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"realpython.com\"</span>, config)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✅ Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(urls)}</span> course posts\"</span>)\n\n    <span class=\"hljs-comment\"># Step 4: Filter for Python tutorials (using metadata!)</span>\n    tutorials = [\n        url <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls \n        <span class=\"hljs-keyword\">if</span> url[<span class=\"hljs-string\">\"status\"</span>] == <span class=\"hljs-string\">\"valid\"</span> <span class=\"hljs-keyword\">and</span> \n        <span class=\"hljs-built_in\">any</span>(keyword <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">str</span>(url[<span class=\"hljs-string\">\"head_data\"</span>]).lower() \n            <span class=\"hljs-keyword\">for</span> keyword <span class=\"hljs-keyword\">in</span> [<span class=\"hljs-string\">\"tutorial\"</span>, <span class=\"hljs-string\">\"guide\"</span>, <span class=\"hljs-string\">\"how to\"</span>])\n    ]\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📚 Filtered to <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(tutorials)}</span> tutorials\"</span>)\n\n    <span class=\"hljs-comment\"># Step 5: Show what we found</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n🎯 Found these tutorials:\"</span>)\n    <span class=\"hljs-keyword\">for</span> tutorial <span class=\"hljs-keyword\">in</span> tutorials[:<span class=\"hljs-number\">5</span>]:  <span class=\"hljs-comment\"># First 5</span>\n        title = tutorial[<span class=\"hljs-string\">\"head_data\"</span>].get(<span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"No title\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  - <span class=\"hljs-subst\">{title}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"    <span class=\"hljs-subst\">{tutorial[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Step 6: Now crawl ONLY these relevant pages</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n🚀 Crawling tutorials...\"</span>)\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        config = CrawlerRunConfig(\n            only_text=<span class=\"hljs-literal\">True</span>,\n            word_count_threshold=<span class=\"hljs-number\">300</span>,  <span class=\"hljs-comment\"># Only substantial articles</span>\n            stream=<span class=\"hljs-literal\">True</span>\n        )\n\n        <span class=\"hljs-comment\"># Extract URLs and crawl them</span>\n        tutorial_urls = [t[<span class=\"hljs-string\">\"url\"</span>] <span class=\"hljs-keyword\">for</span> t <span class=\"hljs-keyword\">in</span> tutorials[:<span class=\"hljs-number\">10</span>]]\n        results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(tutorial_urls, config=config)\n\n        successful = <span class=\"hljs-number\">0</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> results:\n            <span class=\"hljs-keyword\">if</span> result.success:\n                successful += <span class=\"hljs-number\">1</span>\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✓ Crawled: <span class=\"hljs-subst\">{result.url[:<span class=\"hljs-number\">60</span>]}</span>...\"</span>)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n✨ Successfully crawled <span class=\"hljs-subst\">{successful}</span> tutorials!\"</span>)\n\n<span class=\"hljs-comment\"># Run it!</span>\nasyncio.run(smart_blog_crawler())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>What just happened?</strong></p>\n<ol>\n<li>We discovered all blog URLs from the sitemap+cc</li>\n<li>We filtered using metadata (no crawling needed!)</li>\n<li>We crawled only the relevant tutorials</li>\n<li>We saved tons of time and bandwidth</li>\n</ol>\n<p>This is the power of URL seeding - you see everything before you crawl anything.</p>\n<h2 id=\"understanding-the-url-seeder\">Understanding the URL Seeder</h2>\n<p>Now that you've seen the magic, let's understand how it works.</p>\n<h3 id=\"basic-usage\">Basic Usage</h3>\n<p>Creating a URL seeder is simple:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai import AsyncUrlSeeder\n\n<span class=\"hljs-meta\"># Method 1: Manual cleanup</span>\nseeder = AsyncUrlSeeder()\n<span class=\"hljs-keyword\">try</span>:\n    config = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>)\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n<span class=\"hljs-keyword\">finally</span>:\n    <span class=\"hljs-keyword\">await</span> seeder.close()\n\n<span class=\"hljs-meta\"># Method 2: Context manager (recommended)</span>\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncUrlSeeder</span>() <span class=\"hljs-keyword\">as</span> seeder:\n    config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>)\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n    <span class=\"hljs-meta\"># Automatically cleaned up on exit</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The seeder can discover URLs from two powerful sources:</p>\n<h4 id=\"1-sitemaps-fastest\">1. Sitemaps (Fastest)</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Discover from sitemap</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>)\n<span class=\"hljs-attr\">urls</span> = await seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Sitemaps are XML files that websites create specifically to list all their URLs. It's like getting a menu at a restaurant - everything is listed upfront.</p>\n<p><strong>Sitemap Index Support</strong>: For large websites like TechCrunch that use sitemap indexes (a sitemap of sitemaps), the seeder automatically detects and processes all sub-sitemaps in parallel:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\"><span class=\"hljs-comment\">&lt;!-- Example sitemap index --&gt;</span>\n<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">sitemapindex</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">sitemap</span>&gt;</span>\n    <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">loc</span>&gt;</span>https://techcrunch.com/sitemap-1.xml<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">loc</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">sitemap</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">sitemap</span>&gt;</span>\n    <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">loc</span>&gt;</span>https://techcrunch.com/sitemap-2.xml<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">loc</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">sitemap</span>&gt;</span>\n  <span class=\"hljs-comment\">&lt;!-- ... more sitemaps ... --&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">sitemapindex</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The seeder handles this transparently - you'll get all URLs from all sub-sitemaps automatically!</p>\n<h4 id=\"2-common-crawl-most-comprehensive\">2. Common Crawl (Most Comprehensive)</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Discover from Common Crawl</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"cc\"</span>)\n<span class=\"hljs-attr\">urls</span> = await seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Common Crawl is a massive public dataset that regularly crawls the entire web. It's like having access to a pre-built index of the internet.</p>\n<h4 id=\"3-both-sources-maximum-coverage\">3. Both Sources (Maximum Coverage)</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Use both sources</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap+cc\"</span>)\n<span class=\"hljs-attr\">urls</span> = await seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"configuration-magic-seedingconfig\">Configuration Magic: SeedingConfig</h3>\n<p>The <code>SeedingConfig</code> object is your control panel. Here's everything you can configure:</p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Type</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>source</code></td>\n<td>str</td>\n<td>\"sitemap+cc\"</td>\n<td>URL source: \"cc\" (Common Crawl), \"sitemap\", or \"sitemap+cc\"</td>\n</tr>\n<tr>\n<td><code>pattern</code></td>\n<td>str</td>\n<td>\"*\"</td>\n<td>URL pattern filter (e.g., \"<em>/blog/</em>\", \"*.html\")</td>\n</tr>\n<tr>\n<td><code>extract_head</code></td>\n<td>bool</td>\n<td>False</td>\n<td>Extract metadata from page <code>&lt;head&gt;</code></td>\n</tr>\n<tr>\n<td><code>live_check</code></td>\n<td>bool</td>\n<td>False</td>\n<td>Verify URLs are accessible</td>\n</tr>\n<tr>\n<td><code>max_urls</code></td>\n<td>int</td>\n<td>-1</td>\n<td>Maximum URLs to return (-1 = unlimited)</td>\n</tr>\n<tr>\n<td><code>concurrency</code></td>\n<td>int</td>\n<td>10</td>\n<td>Parallel workers for fetching</td>\n</tr>\n<tr>\n<td><code>hits_per_sec</code></td>\n<td>int</td>\n<td>5</td>\n<td>Rate limit for requests</td>\n</tr>\n<tr>\n<td><code>force</code></td>\n<td>bool</td>\n<td>False</td>\n<td>Bypass cache, fetch fresh data</td>\n</tr>\n<tr>\n<td><code>verbose</code></td>\n<td>bool</td>\n<td>False</td>\n<td>Show detailed progress</td>\n</tr>\n<tr>\n<td><code>query</code></td>\n<td>str</td>\n<td>None</td>\n<td>Search query for BM25 scoring</td>\n</tr>\n<tr>\n<td><code>scoring_method</code></td>\n<td>str</td>\n<td>None</td>\n<td>Scoring method (currently \"bm25\")</td>\n</tr>\n<tr>\n<td><code>score_threshold</code></td>\n<td>float</td>\n<td>None</td>\n<td>Minimum score to include URL</td>\n</tr>\n<tr>\n<td><code>filter_nonsense_urls</code></td>\n<td>bool</td>\n<td>True</td>\n<td>Filter out utility URLs (robots.txt, etc.)</td>\n</tr>\n<tr>\n<td><code>cache_ttl_hours</code></td>\n<td>int</td>\n<td>24</td>\n<td>Hours before sitemap cache expires (0 = no TTL)</td>\n</tr>\n<tr>\n<td><code>validate_sitemap_lastmod</code></td>\n<td>bool</td>\n<td>True</td>\n<td>Check sitemap's lastmod and refetch if newer</td>\n</tr>\n</tbody>\n</table>\n<h4 id=\"pattern-matching-examples\">Pattern Matching Examples</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Match all blog posts</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(pattern=<span class=\"hljs-string\">\"*/blog/*\"</span>)\n\n<span class=\"hljs-comment\"># Match only HTML files</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(pattern=<span class=\"hljs-string\">\"*.html\"</span>)\n\n<span class=\"hljs-comment\"># Match product pages</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(pattern=<span class=\"hljs-string\">\"*/product/*\"</span>)\n\n<span class=\"hljs-comment\"># Match everything except admin pages</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(pattern=<span class=\"hljs-string\">\"*\"</span>)\n<span class=\"hljs-comment\"># Then filter: urls = [u for u in urls if \"/admin/\" not in u[\"url\"]]</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"url-validation-live-checking\">URL Validation: Live Checking</h3>\n<p>Sometimes you need to know if URLs are actually accessible. That's where live checking comes in:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    live_check=<span class=\"hljs-literal\">True</span>,  <span class=\"hljs-comment\"># Verify each URL is accessible</span>\n    concurrency=<span class=\"hljs-number\">20</span>    <span class=\"hljs-comment\"># Check 20 URLs in parallel</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Now you can filter by status</span>\nlive_urls = [u <span class=\"hljs-keyword\">for</span> u <span class=\"hljs-keyword\">in</span> urls <span class=\"hljs-keyword\">if</span> u[<span class=\"hljs-string\">\"status\"</span>] == <span class=\"hljs-string\">\"valid\"</span>]\ndead_urls = [u <span class=\"hljs-keyword\">for</span> u <span class=\"hljs-keyword\">in</span> urls <span class=\"hljs-keyword\">if</span> u[<span class=\"hljs-string\">\"status\"</span>] == <span class=\"hljs-string\">\"not_valid\"</span>]\n\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Live URLs: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(live_urls)}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Dead URLs: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(dead_urls)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>When to use live checking:</strong>\n- Before a large crawling operation\n- When working with older sitemaps\n- When data freshness is critical</p>\n<p><strong>When to skip it:</strong>\n- Quick explorations\n- When you trust the source\n- When speed is more important than accuracy</p>\n<h3 id=\"the-power-of-metadata-head-extraction\">The Power of Metadata: Head Extraction</h3>\n<p>This is where URL seeding gets really powerful. Instead of crawling entire pages, you can extract just the metadata:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = SeedingConfig(\n    extract_head=<span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Extract metadata from &lt;head&gt; section</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Now each URL has rich metadata</span>\n<span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls[:<span class=\"hljs-number\">3</span>]:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\nURL: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Title: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'head_data'</span>].get(<span class=\"hljs-string\">'title'</span>)}</span>\"</span>)\n\n    meta = url[<span class=\"hljs-string\">'head_data'</span>].get(<span class=\"hljs-string\">'meta'</span>, {})\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Description: <span class=\"hljs-subst\">{meta.get(<span class=\"hljs-string\">'description'</span>)}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Keywords: <span class=\"hljs-subst\">{meta.get(<span class=\"hljs-string\">'keywords'</span>)}</span>\"</span>)\n\n    <span class=\"hljs-comment\"># Even Open Graph data!</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"OG Image: <span class=\"hljs-subst\">{meta.get(<span class=\"hljs-string\">'og:image'</span>)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"what-can-we-extract\">What Can We Extract?</h4>\n<p>The head extraction gives you a treasure trove of information:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\"><span class=\"hljs-comment\"># Example of extracted head_data</span>\n{\n    <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"10 Python Tips for Beginners\"</span>,\n    <span class=\"hljs-string\">\"charset\"</span>: <span class=\"hljs-string\">\"utf-8\"</span>,\n    <span class=\"hljs-string\">\"lang\"</span>: <span class=\"hljs-string\">\"en\"</span>,\n    <span class=\"hljs-string\">\"meta\"</span>: {\n        <span class=\"hljs-string\">\"description\"</span>: <span class=\"hljs-string\">\"Learn essential Python tips...\"</span>,\n        <span class=\"hljs-string\">\"keywords\"</span>: <span class=\"hljs-string\">\"python, programming, tutorial\"</span>,\n        <span class=\"hljs-string\">\"author\"</span>: <span class=\"hljs-string\">\"Jane Developer\"</span>,\n        <span class=\"hljs-string\">\"viewport\"</span>: <span class=\"hljs-string\">\"width=device-width, initial-scale=1\"</span>,\n\n        <span class=\"hljs-comment\"># Open Graph tags</span>\n        <span class=\"hljs-string\">\"og:title\"</span>: <span class=\"hljs-string\">\"10 Python Tips for Beginners\"</span>,\n        <span class=\"hljs-string\">\"og:description\"</span>: <span class=\"hljs-string\">\"Essential Python tips for new programmers\"</span>,\n        <span class=\"hljs-string\">\"og:image\"</span>: <span class=\"hljs-string\">\"https://example.com/python-tips.jpg\"</span>,\n        <span class=\"hljs-string\">\"og:type\"</span>: <span class=\"hljs-string\">\"article\"</span>,\n\n        <span class=\"hljs-comment\"># Twitter Card tags</span>\n        <span class=\"hljs-string\">\"twitter:card\"</span>: <span class=\"hljs-string\">\"summary_large_image\"</span>,\n        <span class=\"hljs-string\">\"twitter:title\"</span>: <span class=\"hljs-string\">\"10 Python Tips\"</span>,\n\n        <span class=\"hljs-comment\"># Dublin Core metadata</span>\n        <span class=\"hljs-string\">\"dc.creator\"</span>: <span class=\"hljs-string\">\"Jane Developer\"</span>,\n        <span class=\"hljs-string\">\"dc.date\"</span>: <span class=\"hljs-string\">\"2024-01-15\"</span>\n    },\n    <span class=\"hljs-string\">\"link\"</span>: {\n        <span class=\"hljs-string\">\"canonical\"</span>: [{<span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"https://example.com/blog/python-tips\"</span>}],\n        <span class=\"hljs-string\">\"alternate\"</span>: [{<span class=\"hljs-string\">\"href\"</span>: <span class=\"hljs-string\">\"/feed.xml\"</span>, <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"application/rss+xml\"</span>}]\n    },\n    <span class=\"hljs-string\">\"jsonld\"</span>: [\n        {\n            <span class=\"hljs-string\">\"@type\"</span>: <span class=\"hljs-string\">\"Article\"</span>,\n            <span class=\"hljs-string\">\"headline\"</span>: <span class=\"hljs-string\">\"10 Python Tips for Beginners\"</span>,\n            <span class=\"hljs-string\">\"datePublished\"</span>: <span class=\"hljs-string\">\"2024-01-15\"</span>,\n            <span class=\"hljs-string\">\"author\"</span>: {<span class=\"hljs-string\">\"@type\"</span>: <span class=\"hljs-string\">\"Person\"</span>, <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Jane Developer\"</span>}\n        }\n    ]\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This metadata is gold for filtering! You can find exactly what you need without crawling a single page.</p>\n<h3 id=\"smart-url-based-filtering-no-head-extraction\">Smart URL-Based Filtering (No Head Extraction)</h3>\n<p>When <code>extract_head=False</code> but you still provide a query, the seeder uses intelligent URL-based scoring:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Fast filtering based on URL structure alone</span>\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">False</span>,  <span class=\"hljs-comment\"># Don't fetch page metadata</span>\n    query=<span class=\"hljs-string\">\"python tutorial async\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.3</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n\n<span class=\"hljs-comment\"># URLs are scored based on:</span>\n<span class=\"hljs-comment\"># 1. Domain parts matching (e.g., 'python' in python.example.com)</span>\n<span class=\"hljs-comment\"># 2. Path segments (e.g., '/tutorials/python-async/')</span>\n<span class=\"hljs-comment\"># 3. Query parameters (e.g., '?topic=python')</span>\n<span class=\"hljs-comment\"># 4. Fuzzy matching using character n-grams</span>\n\n<span class=\"hljs-comment\"># Example URL scoring:</span>\n<span class=\"hljs-comment\"># https://example.com/tutorials/python/async-guide.html - High score</span>\n<span class=\"hljs-comment\"># https://example.com/blog/javascript-tips.html - Low score</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This approach is much faster than head extraction while still providing intelligent filtering!</p>\n<h3 id=\"understanding-results\">Understanding Results</h3>\n<p>Each URL in the results has this structure:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">{\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://example.com/blog/python-tips.html\"</span>,\n    <span class=\"hljs-string\">\"status\"</span>: <span class=\"hljs-string\">\"valid\"</span>,        <span class=\"hljs-comment\"># \"valid\", \"not_valid\", or \"unknown\"</span>\n    <span class=\"hljs-string\">\"head_data\"</span>: {            <span class=\"hljs-comment\"># Only if extract_head=True</span>\n        <span class=\"hljs-string\">\"title\"</span>: <span class=\"hljs-string\">\"Page Title\"</span>,\n        <span class=\"hljs-string\">\"meta\"</span>: {...},\n        <span class=\"hljs-string\">\"link\"</span>: {...},\n        <span class=\"hljs-string\">\"jsonld\"</span>: [...]\n    },\n    <span class=\"hljs-string\">\"relevance_score\"</span>: 0.85   <span class=\"hljs-comment\"># Only if using BM25 scoring</span>\n}\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Let's see a real example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    live_check=<span class=\"hljs-literal\">True</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"blog.example.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Analyze the results</span>\n<span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls[:<span class=\"hljs-number\">5</span>]:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{<span class=\"hljs-string\">'='</span>*<span class=\"hljs-number\">60</span>}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"URL: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Status: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'status'</span>]}</span>\"</span>)\n\n    <span class=\"hljs-keyword\">if</span> url[<span class=\"hljs-string\">'head_data'</span>]:\n        data = url[<span class=\"hljs-string\">'head_data'</span>]\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Title: <span class=\"hljs-subst\">{data.get(<span class=\"hljs-string\">'title'</span>, <span class=\"hljs-string\">'No title'</span>)}</span>\"</span>)\n\n        <span class=\"hljs-comment\"># Check content type</span>\n        meta = data.get(<span class=\"hljs-string\">'meta'</span>, {})\n        content_type = meta.get(<span class=\"hljs-string\">'og:type'</span>, <span class=\"hljs-string\">'unknown'</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Content Type: <span class=\"hljs-subst\">{content_type}</span>\"</span>)\n\n        <span class=\"hljs-comment\"># Publication date</span>\n        pub_date = <span class=\"hljs-literal\">None</span>\n        <span class=\"hljs-keyword\">for</span> jsonld <span class=\"hljs-keyword\">in</span> data.get(<span class=\"hljs-string\">'jsonld'</span>, []):\n            <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">isinstance</span>(jsonld, <span class=\"hljs-built_in\">dict</span>):\n                pub_date = jsonld.get(<span class=\"hljs-string\">'datePublished'</span>)\n                <span class=\"hljs-keyword\">if</span> pub_date:\n                    <span class=\"hljs-keyword\">break</span>\n\n        <span class=\"hljs-keyword\">if</span> pub_date:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Published: <span class=\"hljs-subst\">{pub_date}</span>\"</span>)\n\n        <span class=\"hljs-comment\"># Word count (if available)</span>\n        word_count = meta.get(<span class=\"hljs-string\">'word_count'</span>)\n        <span class=\"hljs-keyword\">if</span> word_count:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Word Count: <span class=\"hljs-subst\">{word_count}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"smart-filtering-with-bm25-scoring\">Smart Filtering with BM25 Scoring</h2>\n<p>Now for the really cool part - intelligent filtering based on relevance!</p>\n<h3 id=\"introduction-to-relevance-scoring\">Introduction to Relevance Scoring</h3>\n<p>BM25 is a ranking algorithm that scores how relevant a document is to a search query. With URL seeding, we can score URLs based on their metadata <em>before</em> crawling them.</p>\n<p>Think of it like this:\n- Traditional way: Read every book in the library to find ones about Python\n- Smart way: Check the titles and descriptions, score them, read only the most relevant</p>\n<h3 id=\"query-based-discovery\">Query-Based Discovery</h3>\n<p>Here's how to use BM25 scoring:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">config = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,           <span class=\"hljs-comment\"># Required for scoring</span>\n    query=<span class=\"hljs-string\">\"python async tutorial\"</span>,  <span class=\"hljs-comment\"># What we're looking for</span>\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,       <span class=\"hljs-comment\"># Use BM25 algorithm</span>\n    score_threshold=<span class=\"hljs-number\">0.3</span>          <span class=\"hljs-comment\"># Minimum relevance score</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"realpython.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Results are automatically sorted by relevance!</span>\n<span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls[:<span class=\"hljs-number\">5</span>]:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Score: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'relevance_score'</span>]:<span class=\"hljs-number\">.2</span>f}</span> - <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Title: <span class=\"hljs-subst\">{url[<span class=\"hljs-string\">'head_data'</span>][<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"real-examples\">Real Examples</h3>\n<h4 id=\"finding-documentation-pages\">Finding Documentation Pages</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-meta\"># Find API documentation</span>\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=True,\n    query=<span class=\"hljs-string\">\"API reference documentation endpoints\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.5</span>,\n    max_urls=<span class=\"hljs-number\">20</span>\n)\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncUrlSeeder</span>() <span class=\"hljs-keyword\">as</span> seeder:\n    urls</span> = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"docs.example.com\"</span>, config)\n\n<span class=\"hljs-meta\"># The highest scoring URLs will be API docs!</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"discovering-product-pages\">Discovering Product Pages</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-meta\"># Find specific products</span>\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap+cc\"</span>,  <span class=\"hljs-meta\"># Use both sources</span>\n    extract_head=True,\n    query=<span class=\"hljs-string\">\"wireless headphones noise canceling\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.4</span>,\n    pattern=<span class=\"hljs-string\">\"*/product/*\"</span>  <span class=\"hljs-meta\"># Combine with pattern matching</span>\n)\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncUrlSeeder</span>() <span class=\"hljs-keyword\">as</span> seeder:\n    urls</span> = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"shop.example.com\"</span>, config)\n\n<span class=\"hljs-meta\"># Filter further by price (from metadata)</span>\naffordable = [\n    <span class=\"hljs-function\">u <span class=\"hljs-keyword\">for</span> u <span class=\"hljs-keyword\">in</span> urls \n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-title\">float</span>(<span class=\"hljs-params\">u[<span class=\"hljs-string\">'head_data'</span>].<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">'meta'</span>, {}</span>).<span class=\"hljs-title\">get</span>(<span class=\"hljs-params\"><span class=\"hljs-string\">'product:price'</span>, <span class=\"hljs-string\">'0'</span></span>)) &lt; 200\n]\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"filtering-news-articles\">Filtering News Articles</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-meta\"># Find recent news about AI</span>\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=True,\n    query=<span class=\"hljs-string\">\"artificial intelligence machine learning breakthrough\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.35</span>\n)\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncUrlSeeder</span>() <span class=\"hljs-keyword\">as</span> seeder:\n    urls</span> = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"technews.com\"</span>, config)\n\n<span class=\"hljs-meta\"># Filter by date</span>\n<span class=\"hljs-keyword\">from</span> datetime import datetime, timedelta\n\nrecent = []\ncutoff = datetime.now() - timedelta(days=<span class=\"hljs-number\">7</span>)\n\n<span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls:\n    <span class=\"hljs-meta\"># Check JSON-LD for publication date</span>\n    <span class=\"hljs-keyword\">for</span> jsonld <span class=\"hljs-keyword\">in</span> url[<span class=\"hljs-string\">'head_data'</span>].<span class=\"hljs-keyword\">get</span>(<span class=\"hljs-string\">'jsonld'</span>, []):\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-string\">'datePublished'</span> <span class=\"hljs-keyword\">in</span> jsonld:\n            pub_date = datetime.fromisoformat(jsonld[<span class=\"hljs-string\">'datePublished'</span>].replace(<span class=\"hljs-string\">'Z'</span>, <span class=\"hljs-string\">'+00:00'</span>))\n            <span class=\"hljs-keyword\">if</span> pub_date &gt; cutoff:\n                recent.append(url)\n                <span class=\"hljs-keyword\">break</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"complex-query-patterns\">Complex Query Patterns</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Multi-concept queries</span>\nqueries = [\n    <span class=\"hljs-string\">\"python async await concurrency tutorial\"</span>,\n    <span class=\"hljs-string\">\"data science pandas numpy visualization\"</span>,\n    <span class=\"hljs-string\">\"web scraping beautifulsoup selenium automation\"</span>,\n    <span class=\"hljs-string\">\"machine learning tensorflow keras deep learning\"</span>\n]\n\nall_tutorials = []\n\n<span class=\"hljs-keyword\">for</span> query <span class=\"hljs-keyword\">in</span> queries:\n    config = SeedingConfig(\n        source=<span class=\"hljs-string\">\"sitemap\"</span>,\n        extract_head=<span class=\"hljs-literal\">True</span>,\n        query=query,\n        scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n        score_threshold=<span class=\"hljs-number\">0.4</span>,\n        max_urls=<span class=\"hljs-number\">10</span>  <span class=\"hljs-comment\"># Top 10 per topic</span>\n    )\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n        urls = <span class=\"hljs-keyword\">await</span> seeder.urls(<span class=\"hljs-string\">\"learning-platform.com\"</span>, config)\n    all_tutorials.extend(urls)\n\n<span class=\"hljs-comment\"># Remove duplicates while preserving order</span>\nseen = <span class=\"hljs-built_in\">set</span>()\nunique_tutorials = []\n<span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> all_tutorials:\n    <span class=\"hljs-keyword\">if</span> url[<span class=\"hljs-string\">'url'</span>] <span class=\"hljs-keyword\">not</span> <span class=\"hljs-keyword\">in</span> seen:\n        seen.add(url[<span class=\"hljs-string\">'url'</span>])\n        unique_tutorials.append(url)\n\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(unique_tutorials)}</span> unique tutorials across all topics\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"scaling-up-multiple-domains\">Scaling Up: Multiple Domains</h2>\n<p>When you need to discover URLs across multiple websites, URL seeding really shines.</p>\n<h3 id=\"the-many_urls-method\">The <code>many_urls</code> Method</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Discover URLs from multiple domains in parallel</span>\ndomains = [<span class=\"hljs-string\">\"site1.com\"</span>, <span class=\"hljs-string\">\"site2.com\"</span>, <span class=\"hljs-string\">\"site3.com\"</span>]\n\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    query=<span class=\"hljs-string\">\"python tutorial\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.3</span>\n)\n\n<span class=\"hljs-comment\"># Returns a dictionary: {domain: [urls]}</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    results = <span class=\"hljs-keyword\">await</span> seeder.many_urls(domains, config)\n\n<span class=\"hljs-comment\"># Process results</span>\n<span class=\"hljs-keyword\">for</span> domain, urls <span class=\"hljs-keyword\">in</span> results.items():\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{domain}</span>: Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(urls)}</span> relevant URLs\"</span>)\n    <span class=\"hljs-keyword\">if</span> urls:\n        top = urls[<span class=\"hljs-number\">0</span>]  <span class=\"hljs-comment\"># Highest scoring</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Top result: <span class=\"hljs-subst\">{top[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Score: <span class=\"hljs-subst\">{top[<span class=\"hljs-string\">'relevance_score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"cross-domain-examples\">Cross-Domain Examples</h3>\n<h4 id=\"competitor-analysis\">Competitor Analysis</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Analyze content strategies across competitors</span>\ncompetitors = [\n    <span class=\"hljs-string\">\"competitor1.com\"</span>,\n    <span class=\"hljs-string\">\"competitor2.com\"</span>, \n    <span class=\"hljs-string\">\"competitor3.com\"</span>\n]\n\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    pattern=<span class=\"hljs-string\">\"*/blog/*\"</span>,\n    max_urls=<span class=\"hljs-number\">100</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    results = <span class=\"hljs-keyword\">await</span> seeder.many_urls(competitors, config)\n\n<span class=\"hljs-comment\"># Analyze content types</span>\n<span class=\"hljs-keyword\">for</span> domain, urls <span class=\"hljs-keyword\">in</span> results.items():\n    content_types = {}\n\n    <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls:\n        <span class=\"hljs-comment\"># Extract content type from metadata</span>\n        og_type = url[<span class=\"hljs-string\">'head_data'</span>].get(<span class=\"hljs-string\">'meta'</span>, {}).get(<span class=\"hljs-string\">'og:type'</span>, <span class=\"hljs-string\">'unknown'</span>)\n        content_types[og_type] = content_types.get(og_type, <span class=\"hljs-number\">0</span>) + <span class=\"hljs-number\">1</span>\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{domain}</span> content distribution:\"</span>)\n    <span class=\"hljs-keyword\">for</span> ctype, count <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">sorted</span>(content_types.items(), key=<span class=\"hljs-keyword\">lambda</span> x: x[<span class=\"hljs-number\">1</span>], reverse=<span class=\"hljs-literal\">True</span>):\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  <span class=\"hljs-subst\">{ctype}</span>: <span class=\"hljs-subst\">{count}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"industry-research\">Industry Research</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Research Python tutorials across educational sites</span>\neducational_sites = [\n    <span class=\"hljs-string\">\"realpython.com\"</span>,\n    <span class=\"hljs-string\">\"pythontutorial.net\"</span>,\n    <span class=\"hljs-string\">\"learnpython.org\"</span>,\n    <span class=\"hljs-string\">\"python.org\"</span>\n]\n\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"sitemap\"</span>,\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    query=<span class=\"hljs-string\">\"beginner python tutorial basics\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.3</span>,\n    max_urls=<span class=\"hljs-number\">20</span>  <span class=\"hljs-comment\"># Per site</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    results = <span class=\"hljs-keyword\">await</span> seeder.many_urls(educational_sites, config)\n\n<span class=\"hljs-comment\"># Find the best beginner tutorials</span>\nall_tutorials = []\n<span class=\"hljs-keyword\">for</span> domain, urls <span class=\"hljs-keyword\">in</span> results.items():\n    <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls:\n        url[<span class=\"hljs-string\">'domain'</span>] = domain  <span class=\"hljs-comment\"># Add domain info</span>\n        all_tutorials.append(url)\n\n<span class=\"hljs-comment\"># Sort by relevance across all domains</span>\nall_tutorials.sort(key=<span class=\"hljs-keyword\">lambda</span> x: x[<span class=\"hljs-string\">'relevance_score'</span>], reverse=<span class=\"hljs-literal\">True</span>)\n\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Top 10 Python tutorials for beginners across all sites:\"</span>)\n<span class=\"hljs-keyword\">for</span> i, tutorial <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(all_tutorials[:<span class=\"hljs-number\">10</span>], <span class=\"hljs-number\">1</span>):\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{i}</span>. [<span class=\"hljs-subst\">{tutorial[<span class=\"hljs-string\">'relevance_score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>] <span class=\"hljs-subst\">{tutorial[<span class=\"hljs-string\">'head_data'</span>][<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   <span class=\"hljs-subst\">{tutorial[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   From: <span class=\"hljs-subst\">{tutorial[<span class=\"hljs-string\">'domain'</span>]}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h4 id=\"multi-site-monitoring\">Multi-Site Monitoring</h4>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Monitor news about your company across multiple sources</span>\nnews_sites = [\n    <span class=\"hljs-string\">\"techcrunch.com\"</span>,\n    <span class=\"hljs-string\">\"theverge.com\"</span>,\n    <span class=\"hljs-string\">\"wired.com\"</span>,\n    <span class=\"hljs-string\">\"arstechnica.com\"</span>\n]\n\ncompany_name = <span class=\"hljs-string\">\"YourCompany\"</span>\n\nconfig = SeedingConfig(\n    source=<span class=\"hljs-string\">\"cc\"</span>,  <span class=\"hljs-comment\"># Common Crawl for recent content</span>\n    extract_head=<span class=\"hljs-literal\">True</span>,\n    query=<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{company_name}</span> announcement news\"</span>,\n    scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold=<span class=\"hljs-number\">0.5</span>,  <span class=\"hljs-comment\"># High threshold for relevance</span>\n    max_urls=<span class=\"hljs-number\">10</span>\n)\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncUrlSeeder() <span class=\"hljs-keyword\">as</span> seeder:\n    results = <span class=\"hljs-keyword\">await</span> seeder.many_urls(news_sites, config)\n\n<span class=\"hljs-comment\"># Collect all mentions</span>\nmentions = []\n<span class=\"hljs-keyword\">for</span> domain, urls <span class=\"hljs-keyword\">in</span> results.items():\n    mentions.extend(urls)\n\n<span class=\"hljs-keyword\">if</span> mentions:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(mentions)}</span> mentions of <span class=\"hljs-subst\">{company_name}</span>:\"</span>)\n    <span class=\"hljs-keyword\">for</span> mention <span class=\"hljs-keyword\">in</span> mentions:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n- <span class=\"hljs-subst\">{mention[<span class=\"hljs-string\">'head_data'</span>][<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  <span class=\"hljs-subst\">{mention[<span class=\"hljs-string\">'url'</span>]}</span>\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Score: <span class=\"hljs-subst\">{mention[<span class=\"hljs-string\">'relevance_score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>\"</span>)\n<span class=\"hljs-keyword\">else</span>:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"No recent mentions of <span class=\"hljs-subst\">{company_name}</span> found\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"advanced-integration-patterns\">Advanced Integration Patterns</h2>\n<p>Let's put everything together in a real-world example.</p>\n<h3 id=\"building-a-research-assistant\">Building a Research Assistant</h3>\n<p>Here's a complete example that discovers, scores, filters, and crawls intelligently:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> datetime <span class=\"hljs-keyword\">import</span> datetime\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncUrlSeeder, AsyncWebCrawler, SeedingConfig, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">ResearchAssistant</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self</span>):\n        self.seeder = <span class=\"hljs-literal\">None</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__aenter__</span>(<span class=\"hljs-params\">self</span>):\n        self.seeder = AsyncUrlSeeder()\n        <span class=\"hljs-keyword\">await</span> self.seeder.__aenter__()\n        <span class=\"hljs-keyword\">return</span> self\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__aexit__</span>(<span class=\"hljs-params\">self, exc_type, exc_val, exc_tb</span>):\n        <span class=\"hljs-keyword\">if</span> self.seeder:\n            <span class=\"hljs-keyword\">await</span> self.seeder.__aexit__(exc_type, exc_val, exc_tb)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">research_topic</span>(<span class=\"hljs-params\">self, topic, domains, max_articles=<span class=\"hljs-number\">20</span></span>):\n        <span class=\"hljs-string\">\"\"\"Research a topic across multiple domains.\"\"\"</span>\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"🔬 Researching '<span class=\"hljs-subst\">{topic}</span>' across <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(domains)}</span> domains...\"</span>)\n\n        <span class=\"hljs-comment\"># Step 1: Discover relevant URLs</span>\n        config = SeedingConfig(\n            source=<span class=\"hljs-string\">\"sitemap+cc\"</span>,     <span class=\"hljs-comment\"># Maximum coverage</span>\n            extract_head=<span class=\"hljs-literal\">True</span>,       <span class=\"hljs-comment\"># Get metadata</span>\n            query=topic,             <span class=\"hljs-comment\"># Research topic</span>\n            scoring_method=<span class=\"hljs-string\">\"bm25\"</span>,   <span class=\"hljs-comment\"># Smart scoring</span>\n            score_threshold=<span class=\"hljs-number\">0.4</span>,     <span class=\"hljs-comment\"># Quality threshold</span>\n            max_urls=<span class=\"hljs-number\">10</span>,             <span class=\"hljs-comment\"># Per domain</span>\n            concurrency=<span class=\"hljs-number\">20</span>,          <span class=\"hljs-comment\"># Fast discovery</span>\n            verbose=<span class=\"hljs-literal\">True</span>\n        )\n\n        <span class=\"hljs-comment\"># Discover across all domains</span>\n        discoveries = <span class=\"hljs-keyword\">await</span> self.seeder.many_urls(domains, config)\n\n        <span class=\"hljs-comment\"># Step 2: Collect and rank all articles</span>\n        all_articles = []\n        <span class=\"hljs-keyword\">for</span> domain, urls <span class=\"hljs-keyword\">in</span> discoveries.items():\n            <span class=\"hljs-keyword\">for</span> url <span class=\"hljs-keyword\">in</span> urls:\n                url[<span class=\"hljs-string\">'domain'</span>] = domain\n                all_articles.append(url)\n\n        <span class=\"hljs-comment\"># Sort by relevance</span>\n        all_articles.sort(key=<span class=\"hljs-keyword\">lambda</span> x: x[<span class=\"hljs-string\">'relevance_score'</span>], reverse=<span class=\"hljs-literal\">True</span>)\n\n        <span class=\"hljs-comment\"># Take top articles</span>\n        top_articles = all_articles[:max_articles]\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n📊 Found <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(all_articles)}</span> relevant articles\"</span>)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"📌 Selected top <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(top_articles)}</span> for deep analysis\"</span>)\n\n        <span class=\"hljs-comment\"># Step 3: Show what we're about to crawl</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n🎯 Articles to analyze:\"</span>)\n        <span class=\"hljs-keyword\">for</span> i, article <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">enumerate</span>(top_articles[:<span class=\"hljs-number\">5</span>], <span class=\"hljs-number\">1</span>):\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n<span class=\"hljs-subst\">{i}</span>. <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'head_data'</span>][<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Score: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'relevance_score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   Source: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'domain'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"   URL: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'url'</span>][:<span class=\"hljs-number\">60</span>]}</span>...\"</span>)\n\n        <span class=\"hljs-comment\"># Step 4: Crawl the selected articles</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n🚀 Deep crawling <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(top_articles)}</span> articles...\"</span>)\n\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n            config = CrawlerRunConfig(\n                only_text=<span class=\"hljs-literal\">True</span>,\n                word_count_threshold=<span class=\"hljs-number\">200</span>,  <span class=\"hljs-comment\"># Substantial content only</span>\n                stream=<span class=\"hljs-literal\">True</span>\n            )\n\n            <span class=\"hljs-comment\"># Extract URLs and crawl all articles</span>\n            article_urls = [article[<span class=\"hljs-string\">'url'</span>] <span class=\"hljs-keyword\">for</span> article <span class=\"hljs-keyword\">in</span> top_articles]\n            results = []\n            crawl_results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(article_urls, config=config)\n            <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> crawl_results:\n                <span class=\"hljs-keyword\">if</span> result.success:\n                    results.append({\n                        <span class=\"hljs-string\">'url'</span>: result.url,\n                        <span class=\"hljs-string\">'title'</span>: result.metadata.get(<span class=\"hljs-string\">'title'</span>, <span class=\"hljs-string\">'No title'</span>),\n                        <span class=\"hljs-string\">'content'</span>: result.markdown.raw_markdown,\n                        <span class=\"hljs-string\">'domain'</span>: <span class=\"hljs-built_in\">next</span>(a[<span class=\"hljs-string\">'domain'</span>] <span class=\"hljs-keyword\">for</span> a <span class=\"hljs-keyword\">in</span> top_articles <span class=\"hljs-keyword\">if</span> a[<span class=\"hljs-string\">'url'</span>] == result.url),\n                        <span class=\"hljs-string\">'score'</span>: <span class=\"hljs-built_in\">next</span>(a[<span class=\"hljs-string\">'relevance_score'</span>] <span class=\"hljs-keyword\">for</span> a <span class=\"hljs-keyword\">in</span> top_articles <span class=\"hljs-keyword\">if</span> a[<span class=\"hljs-string\">'url'</span>] == result.url)\n                    })\n                    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✓ Crawled: <span class=\"hljs-subst\">{result.url[:<span class=\"hljs-number\">60</span>]}</span>...\"</span>)\n\n        <span class=\"hljs-comment\"># Step 5: Analyze and summarize</span>\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n📝 Analysis complete! Crawled <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(results)}</span> articles\"</span>)\n\n        <span class=\"hljs-keyword\">return</span> self.create_research_summary(topic, results)\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">create_research_summary</span>(<span class=\"hljs-params\">self, topic, articles</span>):\n        <span class=\"hljs-string\">\"\"\"Create a research summary from crawled articles.\"\"\"</span>\n\n        summary = {\n            <span class=\"hljs-string\">'topic'</span>: topic,\n            <span class=\"hljs-string\">'timestamp'</span>: datetime.now().isoformat(),\n            <span class=\"hljs-string\">'total_articles'</span>: <span class=\"hljs-built_in\">len</span>(articles),\n            <span class=\"hljs-string\">'sources'</span>: {}\n        }\n\n        <span class=\"hljs-comment\"># Group by domain</span>\n        <span class=\"hljs-keyword\">for</span> article <span class=\"hljs-keyword\">in</span> articles:\n            domain = article[<span class=\"hljs-string\">'domain'</span>]\n            <span class=\"hljs-keyword\">if</span> domain <span class=\"hljs-keyword\">not</span> <span class=\"hljs-keyword\">in</span> summary[<span class=\"hljs-string\">'sources'</span>]:\n                summary[<span class=\"hljs-string\">'sources'</span>][domain] = []\n\n            summary[<span class=\"hljs-string\">'sources'</span>][domain].append({\n                <span class=\"hljs-string\">'title'</span>: article[<span class=\"hljs-string\">'title'</span>],\n                <span class=\"hljs-string\">'url'</span>: article[<span class=\"hljs-string\">'url'</span>],\n                <span class=\"hljs-string\">'score'</span>: article[<span class=\"hljs-string\">'score'</span>],\n                <span class=\"hljs-string\">'excerpt'</span>: article[<span class=\"hljs-string\">'content'</span>][:<span class=\"hljs-number\">500</span>] + <span class=\"hljs-string\">'...'</span> <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">len</span>(article[<span class=\"hljs-string\">'content'</span>]) &gt; <span class=\"hljs-number\">500</span> <span class=\"hljs-keyword\">else</span> article[<span class=\"hljs-string\">'content'</span>]\n            })\n\n        <span class=\"hljs-keyword\">return</span> summary\n\n<span class=\"hljs-comment\"># Use the research assistant</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> ResearchAssistant() <span class=\"hljs-keyword\">as</span> assistant:\n        <span class=\"hljs-comment\"># Research Python async programming across multiple sources</span>\n        topic = <span class=\"hljs-string\">\"python asyncio best practices performance optimization\"</span>\n        domains = [\n            <span class=\"hljs-string\">\"realpython.com\"</span>,\n            <span class=\"hljs-string\">\"python.org\"</span>,\n            <span class=\"hljs-string\">\"stackoverflow.com\"</span>,\n            <span class=\"hljs-string\">\"medium.com\"</span>\n        ]\n\n        summary = <span class=\"hljs-keyword\">await</span> assistant.research_topic(topic, domains, max_articles=<span class=\"hljs-number\">15</span>)\n\n    <span class=\"hljs-comment\"># Display results</span>\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\n\"</span> + <span class=\"hljs-string\">\"=\"</span>*<span class=\"hljs-number\">60</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"RESEARCH SUMMARY\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"=\"</span>*<span class=\"hljs-number\">60</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Topic: <span class=\"hljs-subst\">{summary[<span class=\"hljs-string\">'topic'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Date: <span class=\"hljs-subst\">{summary[<span class=\"hljs-string\">'timestamp'</span>]}</span>\"</span>)\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total Articles Analyzed: <span class=\"hljs-subst\">{summary[<span class=\"hljs-string\">'total_articles'</span>]}</span>\"</span>)\n\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"\\nKey Findings by Source:\"</span>)\n    <span class=\"hljs-keyword\">for</span> domain, articles <span class=\"hljs-keyword\">in</span> summary[<span class=\"hljs-string\">'sources'</span>].items():\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n📚 <span class=\"hljs-subst\">{domain}</span> (<span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(articles)}</span> articles)\"</span>)\n        <span class=\"hljs-keyword\">for</span> article <span class=\"hljs-keyword\">in</span> articles[:<span class=\"hljs-number\">2</span>]:  <span class=\"hljs-comment\"># Top 2 per domain</span>\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"\\n  Title: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'title'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Relevance: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'score'</span>]:<span class=\"hljs-number\">.2</span>f}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"  Preview: <span class=\"hljs-subst\">{article[<span class=\"hljs-string\">'excerpt'</span>][:<span class=\"hljs-number\">200</span>]}</span>...\"</span>)\n\nasyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"performance-optimization-tips\">Performance Optimization Tips</h3>\n<ol>\n<li>\n<p><strong>Use caching wisely</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># First run - populate cache</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>, extract_head=<span class=\"hljs-literal\">True</span>, force=<span class=\"hljs-literal\">True</span>)\n<span class=\"hljs-attr\">urls</span> = await seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Subsequent runs - use cache (much faster)</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(source=<span class=\"hljs-string\">\"sitemap\"</span>, extract_head=<span class=\"hljs-literal\">True</span>, force=<span class=\"hljs-literal\">False</span>)\n<span class=\"hljs-attr\">urls</span> = await seeder.urls(<span class=\"hljs-string\">\"example.com\"</span>, config)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Optimize concurrency</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># For many small requests (like HEAD checks)</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(concurrency=<span class=\"hljs-number\">50</span>, hits_per_sec=<span class=\"hljs-number\">20</span>)\n\n<span class=\"hljs-comment\"># For fewer large requests (like full head extraction)</span>\n<span class=\"hljs-attr\">config</span> = SeedingConfig(concurrency=<span class=\"hljs-number\">10</span>, hits_per_sec=<span class=\"hljs-number\">5</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Stream large result sets</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># When crawling many URLs</span>\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n    <span class=\"hljs-comment\"># Assuming urls is a list of URL strings</span>\n    crawl_results = <span class=\"hljs-keyword\">await</span> crawler.arun_many(urls, config=config)\n\n    <span class=\"hljs-comment\"># Process as they arrive</span>\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">for</span> result <span class=\"hljs-keyword\">in</span> crawl_results:\n        process_immediately(result)  <span class=\"hljs-comment\"># Don't wait for all</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</li>\n<li>\n<p><strong>Memory protection for large domains</strong></p>\n</li>\n</ol>\n<p>The seeder uses bounded queues to prevent memory issues when processing domains with millions of URLs:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\"><span class=\"hljs-comment\"># Safe for domains with 1M+ URLs</span>\nconfig = SeedingConfig(\n    <span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"cc+sitemap\"</span>,\n    concurrency=50,  <span class=\"hljs-comment\"># Queue size adapts to concurrency</span>\n    max_urls=100000  <span class=\"hljs-comment\"># Process in batches if needed</span>\n)\n\n<span class=\"hljs-comment\"># The seeder automatically manages memory by:</span>\n<span class=\"hljs-comment\"># - Using bounded queues (prevents RAM spikes)</span>\n<span class=\"hljs-comment\"># - Applying backpressure when queue is full</span>\n<span class=\"hljs-comment\"># - Processing URLs as they're discovered</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices-tips\">Best Practices &amp; Tips</h2>\n<h3 id=\"cache-management\">Cache Management</h3>\n<p>The seeder automatically caches results to speed up repeated operations:</p>\n<ul>\n<li><strong>Common Crawl cache</strong>: <code>~/.crawl4ai/seeder_cache/[index]_[domain]_[hash].jsonl</code></li>\n<li><strong>Sitemap cache</strong>: <code>~/.crawl4ai/seeder_cache/sitemap_[domain]_[hash].json</code></li>\n<li><strong>HEAD data cache</strong>: <code>~/.cache/url_seeder/head/[hash].json</code></li>\n</ul>\n<h4 id=\"smart-ttl-cache-for-sitemaps\">Smart TTL Cache for Sitemaps</h4>\n<p>Sitemap caches now include intelligent validation:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Default: 24-hour TTL with lastmod validation</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    cache_ttl_hours<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">24</span>,              <span class=\"hljs-comment\"># Cache expires after 24 hours</span>\n    validate_sitemap_lastmod<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>    <span class=\"hljs-comment\"># Also check if sitemap was updated</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Aggressive caching (1 week, no lastmod check)</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    cache_ttl_hours<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">168</span>,             <span class=\"hljs-comment\"># 7 days</span>\n    validate_sitemap_lastmod<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>   <span class=\"hljs-comment\"># Trust TTL only</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Always validate (no TTL, only lastmod)</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    cache_ttl_hours<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0</span>,               <span class=\"hljs-comment\"># Disable TTL</span>\n    validate_sitemap_lastmod<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>    <span class=\"hljs-comment\"># Refetch if sitemap has newer lastmod</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Always fresh (bypass cache completely)</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    force<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>                       <span class=\"hljs-comment\"># Ignore all caching</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Cache validation priority:</strong>\n1. <code>force=True</code> → Always refetch\n2. Cache doesn't exist → Fetch fresh\n3. <code>validate_sitemap_lastmod=True</code> and sitemap has newer <code>&lt;lastmod&gt;</code> → Refetch\n4. <code>cache_ttl_hours &gt; 0</code> and cache is older than TTL → Refetch\n5. Cache corrupted → Refetch (automatic recovery)\n6. Otherwise → Use cache</p>\n<h3 id=\"pattern-matching-strategies\">Pattern Matching Strategies</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Be specific when possible</span>\ngood_pattern = <span class=\"hljs-string\">\"*/blog/2024/*.html\"</span>  <span class=\"hljs-comment\"># Specific</span>\nbad_pattern = <span class=\"hljs-string\">\"*\"</span>                     <span class=\"hljs-comment\"># Too broad</span>\n\n<span class=\"hljs-comment\"># Combine patterns with metadata filtering</span>\nconfig = SeedingConfig(\n    pattern=<span class=\"hljs-string\">\"*/articles/*\"</span>,\n    extract_head=True\n)\nurls = await seeder.urls(<span class=\"hljs-string\">\"news.com\"</span>, config)\n\n<span class=\"hljs-comment\"># Further filter by publish date, author, category, etc.</span>\nrecent = [u for u in urls if is_recent(u['head_data'])]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"rate-limiting-considerations\">Rate Limiting Considerations</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Be respectful of servers</span>\nconfig = SeedingConfig(\n    hits_per_sec=10,      <span class=\"hljs-comment\"># Max 10 requests per second</span>\n    concurrency=20        <span class=\"hljs-comment\"># But use 20 workers</span>\n)\n\n<span class=\"hljs-comment\"># For your own servers</span>\nconfig = SeedingConfig(\n    hits_per_sec=None,    <span class=\"hljs-comment\"># No limit</span>\n    concurrency=100       <span class=\"hljs-comment\"># Go fast</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"quick-reference\">Quick Reference</h2>\n<h3 id=\"common-patterns\">Common Patterns</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Blog post discovery</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    pattern<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"*/blog/*\"</span>,\n    extract_head<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"your topic\"</span>,\n    scoring_method<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"bm25\"</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># E-commerce product discovery</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap+cc\"</span>,\n    pattern<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"*/product/*\"</span>,\n    extract_head<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    live_check<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Documentation search</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    pattern<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"*/docs/*\"</span>,\n    extract_head<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"API reference\"</span>,\n    scoring_method<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"bm25\"</span>,\n    score_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.5</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># News monitoring</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"cc\"</span>,\n    extract_head<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"company name\"</span>,\n    scoring_method<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"bm25\"</span>,\n    max_urls<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">50</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"troubleshooting-guide\">Troubleshooting Guide</h3>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Issue</th>\n<th>Solution</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No URLs found</td>\n<td>Try <code>source=\"cc+sitemap\"</code>, check domain spelling</td>\n</tr>\n<tr>\n<td>Slow discovery</td>\n<td>Reduce <code>concurrency</code>, add <code>hits_per_sec</code> limit</td>\n</tr>\n<tr>\n<td>Missing metadata</td>\n<td>Ensure <code>extract_head=True</code></td>\n</tr>\n<tr>\n<td>Low relevance scores</td>\n<td>Refine query, lower <code>score_threshold</code></td>\n</tr>\n<tr>\n<td>Rate limit errors</td>\n<td>Reduce <code>hits_per_sec</code> and <code>concurrency</code></td>\n</tr>\n<tr>\n<td>Memory issues with large sites</td>\n<td>Use <code>max_urls</code> to limit results, reduce <code>concurrency</code></td>\n</tr>\n<tr>\n<td>Connection not closed</td>\n<td>Use context manager or call <code>await seeder.close()</code></td>\n</tr>\n<tr>\n<td>Stale/outdated URLs</td>\n<td>Set <code>cache_ttl_hours=0</code> or use <code>force=True</code></td>\n</tr>\n<tr>\n<td>Cache not updating</td>\n<td>Check <code>validate_sitemap_lastmod=True</code>, or use <code>force=True</code></td>\n</tr>\n<tr>\n<td>Incomplete URL list</td>\n<td>Delete cache file and refetch, or use <code>force=True</code></td>\n</tr>\n</tbody>\n</table>\n<h3 id=\"performance-benchmarks\">Performance Benchmarks</h3>\n<p>Typical performance on a standard connection:</p>\n<ul>\n<li><strong>Sitemap discovery</strong>: 100-1,000 URLs/second</li>\n<li><strong>Common Crawl discovery</strong>: 50-500 URLs/second  </li>\n<li><strong>HEAD checking</strong>: 10-50 URLs/second</li>\n<li><strong>Head extraction</strong>: 5-20 URLs/second</li>\n<li><strong>BM25 scoring</strong>: 10,000+ URLs/second</li>\n</ul>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>URL seeding transforms web crawling from a blind expedition into a surgical strike. By discovering and analyzing URLs before crawling, you can:</p>\n<ul>\n<li>Save hours of crawling time</li>\n<li>Reduce bandwidth usage by 90%+</li>\n<li>Find exactly what you need</li>\n<li>Scale across multiple domains effortlessly</li>\n</ul>\n<p>Whether you're building a research tool, monitoring competitors, or creating a content aggregator, URL seeding gives you the intelligence to crawl smarter, not harder.</p>\n<h3 id=\"smart-url-filtering\">Smart URL Filtering</h3>\n<p>The seeder automatically filters out nonsense URLs that aren't useful for content crawling:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Enabled by default</span>\nconfig <span class=\"hljs-punctuation\">=</span> SeedingConfig<span class=\"hljs-punctuation\">(</span>\n    source<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"sitemap\"</span>,\n    filter_nonsense_urls<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>  <span class=\"hljs-comment\"># Default: True</span>\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># URLs that get filtered:</span>\n<span class=\"hljs-comment\"># - robots.txt, sitemap.xml, ads.txt</span>\n<span class=\"hljs-comment\"># - API endpoints (/api/, /v1/, .json)</span>\n<span class=\"hljs-comment\"># - Media files (.jpg, .mp4, .pdf)</span>\n<span class=\"hljs-comment\"># - Archives (.zip, .tar.gz)</span>\n<span class=\"hljs-comment\"># - Source code (.js, .css)</span>\n<span class=\"hljs-comment\"># - Admin/login pages</span>\n<span class=\"hljs-comment\"># - And many more...</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>To disable filtering (not recommended):</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">config = SeedingConfig(\n    <span class=\"hljs-built_in\">source</span>=<span class=\"hljs-string\">\"sitemap\"</span>,\n    filter_nonsense_urls=False  <span class=\"hljs-comment\"># Include ALL URLs</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"key-features-summary\">Key Features Summary</h3>\n<ol>\n<li><strong>Parallel Sitemap Index Processing</strong>: Automatically detects and processes sitemap indexes in parallel</li>\n<li><strong>Memory Protection</strong>: Bounded queues prevent RAM issues with large domains (1M+ URLs)</li>\n<li><strong>Context Manager Support</strong>: Automatic cleanup with <code>async with</code> statement</li>\n<li><strong>URL-Based Scoring</strong>: Smart filtering even without head extraction</li>\n<li><strong>Smart URL Filtering</strong>: Automatically excludes utility/nonsense URLs</li>\n<li><strong>Smart TTL Cache</strong>: Sitemap caches with TTL expiry and lastmod validation</li>\n<li><strong>Automatic Cache Recovery</strong>: Corrupted or incomplete caches are automatically refreshed</li>\n</ol>\n<p>Now go forth and seed intelligently!</p>\n<h2 id=\"need-more-coverage\">Need More Coverage?</h2>\n<p>If you need to discover URLs across an entire domain — including subdomains, hidden services, and pages not listed in any sitemap — check out <a href=\"../domain-mapping/\">Domain Mapping</a>. It combines 8 discovery sources (including Certificate Transparency, Wayback Machine, path probing, and soft-404 detection) to find everything under a domain.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/extraction/chunking/",
    "title": "Chunking - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"chunking-strategies\">Chunking Strategies</h1>\n<p>Chunking strategies are critical for dividing large texts into manageable parts, enabling effective content processing and extraction. These strategies are foundational in cosine similarity-based extraction techniques, which allow users to retrieve only the most relevant chunks of content for a given query. Additionally, they facilitate direct integration into RAG (Retrieval-Augmented Generation) systems for structured and scalable workflows.</p>\n<h3 id=\"why-use-chunking\">Why Use Chunking?</h3>\n<p>1. <strong>Cosine Similarity and Query Relevance</strong>: Prepares chunks for semantic similarity analysis.\n2. <strong>RAG System Integration</strong>: Seamlessly processes and stores chunks for retrieval.\n3. <strong>Structured Processing</strong>: Allows for diverse segmentation methods, such as sentence-based, topic-based, or windowed approaches.</p>\n<h3 id=\"methods-of-chunking\">Methods of Chunking</h3>\n<h4 id=\"1-regex-based-chunking\">1. Regex-Based Chunking</h4>\n<p>Splits text based on regular expression patterns, useful for coarse segmentation.</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">RegexChunking</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, patterns=<span class=\"hljs-literal\">None</span></span>):\n        self.patterns = patterns <span class=\"hljs-keyword\">or</span> [<span class=\"hljs-string\">r'\\n\\n'</span>]  <span class=\"hljs-comment\"># Default pattern for paragraphs</span>\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">chunk</span>(<span class=\"hljs-params\">self, text</span>):\n        paragraphs = [text]\n        <span class=\"hljs-keyword\">for</span> pattern <span class=\"hljs-keyword\">in</span> self.patterns:\n            paragraphs = [seg <span class=\"hljs-keyword\">for</span> p <span class=\"hljs-keyword\">in</span> paragraphs <span class=\"hljs-keyword\">for</span> seg <span class=\"hljs-keyword\">in</span> re.split(pattern, p)]\n        <span class=\"hljs-keyword\">return</span> paragraphs\n\n<span class=\"hljs-comment\"># Example Usage</span>\ntext = <span class=\"hljs-string\">\"\"\"This is the first paragraph.\n\nThis is the second paragraph.\"\"\"</span>\nchunker = RegexChunking()\n<span class=\"hljs-built_in\">print</span>(chunker.chunk(text))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h4 id=\"2-sentence-based-chunking\">2. Sentence-Based Chunking</h4>\n<p>Divides text into sentences using NLP tools, ideal for extracting meaningful statements.</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> nltk.tokenize <span class=\"hljs-keyword\">import</span> sent_tokenize\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">NlpSentenceChunking</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">chunk</span>(<span class=\"hljs-params\">self, text</span>):\n        sentences = sent_tokenize(text)\n        <span class=\"hljs-keyword\">return</span> [sentence.strip() <span class=\"hljs-keyword\">for</span> sentence <span class=\"hljs-keyword\">in</span> sentences]\n\n<span class=\"hljs-comment\"># Example Usage</span>\ntext = <span class=\"hljs-string\">\"This is sentence one. This is sentence two.\"</span>\nchunker = NlpSentenceChunking()\n<span class=\"hljs-built_in\">print</span>(chunker.chunk(text))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h4 id=\"3-topic-based-segmentation\">3. Topic-Based Segmentation</h4>\n<p>Uses algorithms like TextTiling to create topic-coherent chunks.</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> nltk.tokenize <span class=\"hljs-keyword\">import</span> TextTilingTokenizer\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">TopicSegmentationChunking</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self</span>):\n        self.tokenizer = TextTilingTokenizer()\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">chunk</span>(<span class=\"hljs-params\">self, text</span>):\n        <span class=\"hljs-keyword\">return</span> self.tokenizer.tokenize(text)\n\n<span class=\"hljs-comment\"># Example Usage</span>\ntext = <span class=\"hljs-string\">\"\"\"This is an introduction.\nThis is a detailed discussion on the topic.\"\"\"</span>\nchunker = TopicSegmentationChunking()\n<span class=\"hljs-built_in\">print</span>(chunker.chunk(text))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h4 id=\"4-fixed-length-word-chunking\">4. Fixed-Length Word Chunking</h4>\n<p>Segments text into chunks of a fixed word count.</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">FixedLengthWordChunking</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, chunk_size=<span class=\"hljs-number\">100</span></span>):\n        self.chunk_size = chunk_size\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">chunk</span>(<span class=\"hljs-params\">self, text</span>):\n        words = text.split()\n        <span class=\"hljs-keyword\">return</span> [<span class=\"hljs-string\">' '</span>.join(words[i:i + self.chunk_size]) <span class=\"hljs-keyword\">for</span> i <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">0</span>, <span class=\"hljs-built_in\">len</span>(words), self.chunk_size)]\n\n<span class=\"hljs-comment\"># Example Usage</span>\ntext = <span class=\"hljs-string\">\"This is a long text with many words to be chunked into fixed sizes.\"</span>\nchunker = FixedLengthWordChunking(chunk_size=<span class=\"hljs-number\">5</span>)\n<span class=\"hljs-built_in\">print</span>(chunker.chunk(text))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h4 id=\"5-sliding-window-chunking\">5. Sliding Window Chunking</h4>\n<p>Generates overlapping chunks for better contextual coherence.</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">SlidingWindowChunking</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, window_size=<span class=\"hljs-number\">100</span>, step=<span class=\"hljs-number\">50</span></span>):\n        self.window_size = window_size\n        self.step = step\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">chunk</span>(<span class=\"hljs-params\">self, text</span>):\n        words = text.split()\n        chunks = []\n        <span class=\"hljs-keyword\">for</span> i <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-number\">0</span>, <span class=\"hljs-built_in\">len</span>(words) - self.window_size + <span class=\"hljs-number\">1</span>, self.step):\n            chunks.append(<span class=\"hljs-string\">' '</span>.join(words[i:i + self.window_size]))\n        <span class=\"hljs-keyword\">return</span> chunks\n\n<span class=\"hljs-comment\"># Example Usage</span>\ntext = <span class=\"hljs-string\">\"This is a long text to demonstrate sliding window chunking.\"</span>\nchunker = SlidingWindowChunking(window_size=<span class=\"hljs-number\">5</span>, step=<span class=\"hljs-number\">2</span>)\n<span class=\"hljs-built_in\">print</span>(chunker.chunk(text))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"combining-chunking-with-cosine-similarity\">Combining Chunking with Cosine Similarity</h3>\n<p>To enhance the relevance of extracted content, chunking strategies can be paired with cosine similarity techniques. Here’s an example workflow:</p>\n<p><strong>Code Example</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> sklearn.feature_extraction.text <span class=\"hljs-keyword\">import</span> TfidfVectorizer\n<span class=\"hljs-keyword\">from</span> sklearn.metrics.pairwise <span class=\"hljs-keyword\">import</span> cosine_similarity\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CosineSimilarityExtractor</span>:\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\">self, query</span>):\n        self.query = query\n        self.vectorizer = TfidfVectorizer()\n\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">find_relevant_chunks</span>(<span class=\"hljs-params\">self, chunks</span>):\n        vectors = self.vectorizer.fit_transform([self.query] + chunks)\n        similarities = cosine_similarity(vectors[<span class=\"hljs-number\">0</span>:<span class=\"hljs-number\">1</span>], vectors[<span class=\"hljs-number\">1</span>:]).flatten()\n        <span class=\"hljs-keyword\">return</span> [(chunks[i], similarities[i]) <span class=\"hljs-keyword\">for</span> i <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">range</span>(<span class=\"hljs-built_in\">len</span>(chunks))]\n\n<span class=\"hljs-comment\"># Example Workflow</span>\ntext = <span class=\"hljs-string\">\"\"\"This is a sample document. It has multiple sentences. \nWe are testing chunking and similarity.\"\"\"</span>\n\nchunker = SlidingWindowChunking(window_size=<span class=\"hljs-number\">5</span>, step=<span class=\"hljs-number\">3</span>)\nchunks = chunker.chunk(text)\nquery = <span class=\"hljs-string\">\"testing chunking\"</span>\nextractor = CosineSimilarityExtractor(query)\nrelevant_chunks = extractor.find_relevant_chunks(chunks)\n\n<span class=\"hljs-built_in\">print</span>(relevant_chunks)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/extraction/clustring-strategies/",
    "title": "Clustering Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"cosine-strategy\">Cosine Strategy</h1>\n<p>The Cosine Strategy in Crawl4AI uses similarity-based clustering to identify and extract relevant content sections from web pages. This strategy is particularly useful when you need to find and extract content based on semantic similarity rather than structural patterns.</p>\n<h2 id=\"how-it-works\">How It Works</h2>\n<p>The Cosine Strategy:\n1. Breaks down page content into meaningful chunks\n2. Converts text into vector representations\n3. Calculates similarity between chunks\n4. Clusters similar content together\n5. Ranks and filters content based on relevance</p>\n<h2 id=\"basic-usage\">Basic Usage</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-csharp\"><span class=\"hljs-keyword\">from</span> crawl4ai import CosineStrategy\n\nstrategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"product reviews\"</span>,    <span class=\"hljs-meta\"># Target content type</span>\n    word_count_threshold=<span class=\"hljs-number\">10</span>,             <span class=\"hljs-meta\"># Minimum words per cluster</span>\n    sim_threshold=<span class=\"hljs-number\">0.3</span>                    <span class=\"hljs-meta\"># Similarity threshold</span>\n)\n\n<span class=\"hljs-function\"><span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> <span class=\"hljs-title\">AsyncWebCrawler</span>() <span class=\"hljs-keyword\">as</span> crawler:\n    result</span> = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com/reviews\"</span>,\n        extraction_strategy=strategy\n    )\n\n    content = result.extracted_content\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"configuration-options\">Configuration Options</h2>\n<h3 id=\"core-parameters\">Core Parameters</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">CosineStrategy(\n    <span class=\"hljs-comment\"># Content Filtering</span>\n    semantic_filter: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-literal\">None</span>,       <span class=\"hljs-comment\"># Keywords/topic for content filtering</span>\n    word_count_threshold: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">10</span>,    <span class=\"hljs-comment\"># Minimum words per cluster</span>\n    sim_threshold: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.3</span>,        <span class=\"hljs-comment\"># Similarity threshold (0.0 to 1.0)</span>\n\n    <span class=\"hljs-comment\"># Clustering Parameters</span>\n    max_dist: <span class=\"hljs-built_in\">float</span> = <span class=\"hljs-number\">0.2</span>,            <span class=\"hljs-comment\"># Maximum distance for clustering</span>\n    linkage_method: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">'ward'</span>,      <span class=\"hljs-comment\"># Clustering linkage method</span>\n    top_k: <span class=\"hljs-built_in\">int</span> = <span class=\"hljs-number\">3</span>,                   <span class=\"hljs-comment\"># Number of top categories to extract</span>\n\n    <span class=\"hljs-comment\"># Model Configuration</span>\n    model_name: <span class=\"hljs-built_in\">str</span> = <span class=\"hljs-string\">'sentence-transformers/all-MiniLM-L6-v2'</span>,  <span class=\"hljs-comment\"># Embedding model</span>\n\n    verbose: <span class=\"hljs-built_in\">bool</span> = <span class=\"hljs-literal\">False</span>             <span class=\"hljs-comment\"># Enable logging</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"parameter-details\">Parameter Details</h3>\n<p>1. <strong>semantic_filter</strong>\n   - Sets the target topic or content type\n   - Use keywords relevant to your desired content\n   - Example: \"technical specifications\", \"user reviews\", \"pricing information\"</p>\n<p>2. <strong>sim_threshold</strong>\n   - Controls how similar content must be to be grouped together\n   - Higher values (e.g., 0.8) mean stricter matching\n   - Lower values (e.g., 0.3) allow more variation\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Strict matching</span>\n<span class=\"hljs-attr\">strategy</span> = CosineStrategy(sim_threshold=<span class=\"hljs-number\">0.8</span>)\n\n<span class=\"hljs-comment\"># Loose matching</span>\n<span class=\"hljs-attr\">strategy</span> = CosineStrategy(sim_threshold=<span class=\"hljs-number\">0.3</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>3. <strong>word_count_threshold</strong>\n   - Filters out short content blocks\n   - Helps eliminate noise and irrelevant content\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Only consider substantial paragraphs</span>\n<span class=\"hljs-attr\">strategy</span> = CosineStrategy(word_count_threshold=<span class=\"hljs-number\">50</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>4. <strong>top_k</strong>\n   - Number of top content clusters to return\n   - Higher values return more diverse content\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Get top 5 most relevant content clusters</span>\n<span class=\"hljs-attr\">strategy</span> = CosineStrategy(top_k=<span class=\"hljs-number\">5</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"use-cases\">Use Cases</h2>\n<h3 id=\"1-article-content-extraction\">1. Article Content Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">strategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"main article content\"</span>,\n    word_count_threshold=100,  <span class=\"hljs-comment\"># Longer blocks for articles</span>\n    top_k=1                   <span class=\"hljs-comment\"># Usually want single main content</span>\n)\n\nresult = await crawler.arun(\n    url=<span class=\"hljs-string\">\"https://example.com/blog/post\"</span>,\n    extraction_strategy=strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-product-review-analysis\">2. Product Review Analysis</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">strategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"customer reviews and ratings\"</span>,\n    word_count_threshold=20,   <span class=\"hljs-comment\"># Reviews can be shorter</span>\n    top_k=10,                 <span class=\"hljs-comment\"># Get multiple reviews</span>\n    sim_threshold=0.4         <span class=\"hljs-comment\"># Allow variety in review content</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-technical-documentation\">3. Technical Documentation</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">strategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"technical specifications documentation\"</span>,\n    word_count_threshold=30,\n    sim_threshold=0.6,        <span class=\"hljs-comment\"># Stricter matching for technical content</span>\n    max_dist=0.3             <span class=\"hljs-comment\"># Allow related technical sections</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"advanced-features\">Advanced Features</h2>\n<h3 id=\"custom-clustering\">Custom Clustering</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">strategy = CosineStrategy(\n    linkage_method=<span class=\"hljs-string\">'complete'</span>,  <span class=\"hljs-comment\"># Alternative clustering method</span>\n    max_dist=0.4,              <span class=\"hljs-comment\"># Larger clusters</span>\n    model_name=<span class=\"hljs-string\">'sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2'</span>  <span class=\"hljs-comment\"># Multilingual support</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"content-filtering-pipeline\">Content Filtering Pipeline</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\">strategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"pricing plans features\"</span>,\n    word_count_threshold=<span class=\"hljs-number\">15</span>,\n    sim_threshold=<span class=\"hljs-number\">0.5</span>,\n    top_k=<span class=\"hljs-number\">3</span>\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_pricing_features</span>(<span class=\"hljs-params\">url: <span class=\"hljs-built_in\">str</span></span>):\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=url,\n            extraction_strategy=strategy\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            content = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">return</span> {\n                <span class=\"hljs-string\">'pricing_features'</span>: content,\n                <span class=\"hljs-string\">'clusters'</span>: <span class=\"hljs-built_in\">len</span>(content),\n                <span class=\"hljs-string\">'similarity_scores'</span>: [item[<span class=\"hljs-string\">'score'</span>] <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> content]\n            }\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"best-practices\">Best Practices</h2>\n<p>1. <strong>Adjust Thresholds Iteratively</strong>\n   - Start with default values\n   - Adjust based on results\n   - Monitor clustering quality</p>\n<p>2. <strong>Choose Appropriate Word Count Thresholds</strong>\n   - Higher for articles (100+)\n   - Lower for reviews/comments (20+)\n   - Medium for product descriptions (50+)</p>\n<p>3. <strong>Optimize Performance</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">strategy <span class=\"hljs-punctuation\">=</span> CosineStrategy<span class=\"hljs-punctuation\">(</span>\n    word_count_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">10</span>,  <span class=\"hljs-comment\"># Filter early</span>\n    top_k<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">5</span>,                 <span class=\"hljs-comment\"># Limit results</span>\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>             <span class=\"hljs-comment\"># Monitor performance</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p>4. <strong>Handle Different Content Types</strong>\n   </p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># For mixed content pages</span>\nstrategy = CosineStrategy(\n    semantic_filter=<span class=\"hljs-string\">\"product features\"</span>,\n    sim_threshold=0.4,      <span class=\"hljs-comment\"># More flexible matching</span>\n    max_dist=0.3,          <span class=\"hljs-comment\"># Larger clusters</span>\n    top_k=3                <span class=\"hljs-comment\"># Multiple relevant sections</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"error-handling\">Error Handling</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">try</span>:\n    result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n        url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n        extraction_strategy=strategy\n    )\n\n    <span class=\"hljs-keyword\">if</span> result.success:\n        content = json.loads(result.extracted_content)\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> content:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"No relevant content found\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extraction failed: <span class=\"hljs-subst\">{result.error_message}</span>\"</span>)\n\n<span class=\"hljs-keyword\">except</span> Exception <span class=\"hljs-keyword\">as</span> e:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Error during extraction: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">str</span>(e)}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The Cosine Strategy is particularly effective when:\n- Content structure is inconsistent\n- You need semantic understanding\n- You want to find similar content blocks\n- Structure-based extraction (CSS/XPath) isn't reliable</p>\n<p>It works well with other strategies and can be used as a pre-processing step for LLM-based extraction.</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/extraction/llm-strategies/",
    "title": "LLM Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"extracting-json-llm\">Extracting JSON (LLM)</h1>\n<p>In some cases, you need to extract <strong>complex or unstructured</strong> information from a webpage that a simple CSS/XPath schema cannot easily parse. Or you want <strong>AI</strong>-driven insights, classification, or summarization. For these scenarios, Crawl4AI provides an <strong>LLM-based extraction strategy</strong> that:</p>\n<ol>\n<li>Works with <strong>any</strong> large language model supported by <a href=\"https://github.com/BerriAI/litellm\">LiteLLM</a> (Ollama, OpenAI, Claude, and more).  </li>\n<li>Automatically splits content into chunks (if desired) to handle token limits, then combines results.  </li>\n<li>Lets you define a <strong>schema</strong> (like a Pydantic model) or a simpler “block” extraction approach.</li>\n</ol>\n<p><strong>Important</strong>: LLM-based extraction can be slower and costlier than schema-based approaches. If your page data is highly structured, consider using <a href=\"../no-llm-strategies/\"><code>JsonCssExtractionStrategy</code></a> or <a href=\"../no-llm-strategies/\"><code>JsonXPathExtractionStrategy</code></a> first. But if you need AI to interpret or reorganize content, read on!</p>\n<hr>\n<h2 id=\"1-why-use-an-llm\">1. Why Use an LLM?</h2>\n<ul>\n<li><strong>Complex Reasoning</strong>: If the site’s data is unstructured, scattered, or full of natural language context.  </li>\n<li><strong>Semantic Extraction</strong>: Summaries, knowledge graphs, or relational data that require comprehension.  </li>\n<li><strong>Flexible</strong>: You can pass instructions to the model to do more advanced transformations or classification.</li>\n</ul>\n<hr>\n<h2 id=\"2-provider-agnostic-via-litellm\">2. Provider-Agnostic via LiteLLM</h2>\n<p>You can use LLMConfig, to quickly configure multiple variations of LLMs and experiment with them to find the optimal one for your use case. You can read more about LLMConfig <a href=\"/api/parameters\">here</a>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-attr\">llm_config</span> = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>, api_token=os.getenv(<span class=\"hljs-string\">\"OPENAI_API_KEY\"</span>))\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Crawl4AI uses a “provider string” (e.g., <code>\"openai/gpt-4o\"</code>, <code>\"ollama/llama2.0\"</code>, <code>\"aws/titan\"</code>) to identify your LLM. <strong>Any</strong> model that LiteLLM supports is fair game. You just provide:</p>\n<ul>\n<li><strong><code>provider</code></strong>: The <code>&lt;provider&gt;/&lt;model_name&gt;</code> identifier (e.g., <code>\"openai/gpt-4\"</code>, <code>\"ollama/llama2\"</code>, <code>\"huggingface/google-flan\"</code>, etc.).  </li>\n<li><strong><code>api_token</code></strong>: If needed (for OpenAI, HuggingFace, etc.); local models or Ollama might not require it.  </li>\n<li><strong><code>base_url</code></strong> (optional): If your provider has a custom endpoint.  </li>\n</ul>\n<p>This means you <strong>aren’t locked</strong> into a single LLM vendor. Switch or experiment easily.</p>\n<hr>\n<h2 id=\"3-how-llm-extraction-works\">3. How LLM Extraction Works</h2>\n<h3 id=\"31-flow\">3.1 Flow</h3>\n<p>1. <strong>Chunking</strong> (optional): The HTML or markdown is split into smaller segments if it’s very long (based on <code>chunk_token_threshold</code>, overlap, etc.).<br>\n2. <strong>Prompt Construction</strong>: For each chunk, the library forms a prompt that includes your <strong><code>instruction</code></strong> (and possibly schema or examples).<br>\n3. <strong>LLM Inference</strong>: Each chunk is sent to the model in parallel or sequentially (depending on your concurrency).<br>\n4. <strong>Combining</strong>: The results from each chunk are merged and parsed into JSON.</p>\n<h3 id=\"32-extraction_type\">3.2 <code>extraction_type</code></h3>\n<ul>\n<li><strong><code>\"schema\"</code></strong>: The model tries to return JSON conforming to your Pydantic-based schema.  </li>\n<li><strong><code>\"block\"</code></strong>: The model returns freeform text, or smaller JSON structures, which the library collects.  </li>\n</ul>\n<p>For structured data, <code>\"schema\"</code> is recommended. You provide <code>schema=YourPydanticModel.model_json_schema()</code>.</p>\n<hr>\n<h2 id=\"4-key-parameters\">4. Key Parameters</h2>\n<p>Below is an overview of important LLM extraction parameters. All are typically set inside <code>LLMExtractionStrategy(...)</code>. You then put that strategy in your <code>CrawlerRunConfig(..., extraction_strategy=...)</code>.</p>\n<ol>\n<li><strong><code>llm_config</code></strong> (LLMConfig): e.g., <code>\"openai/gpt-4\"</code>, <code>\"ollama/llama2\"</code>.\n2. <strong><code>schema</code></strong> (dict): A JSON schema describing the fields you want. Usually generated by <code>YourModel.model_json_schema()</code>.<br>\n3. <strong><code>extraction_type</code></strong> (str): <code>\"schema\"</code> or <code>\"block\"</code>.<br>\n4. <strong><code>instruction</code></strong> (str): Prompt text telling the LLM what you want extracted. E.g., “Extract these fields as a JSON array.”<br>\n5. <strong><code>chunk_token_threshold</code></strong> (int): Maximum tokens per chunk. If your content is huge, you can break it up for the LLM.<br>\n6. <strong><code>overlap_rate</code></strong> (float): Overlap ratio between adjacent chunks. E.g., <code>0.1</code> means 10% of each chunk is repeated to preserve context continuity.<br>\n7. <strong><code>apply_chunking</code></strong> (bool): Set <code>True</code> to chunk automatically. If you want a single pass, set <code>False</code>.<br>\n8. <strong><code>input_format</code></strong> (str): Determines <strong>which</strong> crawler result is passed to the LLM. Options include:  </li>\n<li><code>\"markdown\"</code>: The raw markdown (default).  </li>\n<li><code>\"fit_markdown\"</code>: The filtered “fit” markdown if you used a content filter.  </li>\n<li><code>\"html\"</code>: The cleaned or raw HTML.<br>\n9. <strong><code>extra_args</code></strong> (dict): Additional LLM parameters like <code>temperature</code>, <code>max_tokens</code>, <code>top_p</code>, etc.<br>\n10. <strong><code>show_usage()</code></strong>: A method you can call to print out usage info (token usage per chunk, total cost if known).  </li>\n</ol>\n<p><strong>Example</strong>:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\">extraction_strategy <span class=\"hljs-punctuation\">=</span> LLMExtractionStrategy<span class=\"hljs-punctuation\">(</span>\n    llm_config <span class=\"hljs-punctuation\">=</span> LLMConfig<span class=\"hljs-punctuation\">(</span>provider<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"openai/gpt-4\"</span>, api_token<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"YOUR_OPENAI_KEY\"</span><span class=\"hljs-punctuation\">)</span>,\n    <span class=\"hljs-keyword\">schema</span><span class=\"hljs-punctuation\">=</span>MyModel.model_json_schema<span class=\"hljs-punctuation\">(</span><span class=\"hljs-punctuation\">)</span>,\n    extraction_type<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"schema\"</span>,\n    instruction<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Extract a list of items from the text with 'name' and 'price' fields.\"</span>,\n    chunk_token_threshold<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">1200</span>,\n    overlap_rate<span class=\"hljs-punctuation\">=</span><span class=\"hljs-number\">0.1</span>,\n    apply_chunking<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>,\n    input_format<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"html\"</span>,\n    extra_args<span class=\"hljs-punctuation\">=</span><span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"temperature\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">0.1</span>, <span class=\"hljs-string\">\"max_tokens\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-number\">1000</span><span class=\"hljs-punctuation\">}</span>,\n    verbose<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">True</span>\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"5-putting-it-in-crawlerrunconfig\">5. Putting It in <code>CrawlerRunConfig</code></h2>\n<p><strong>Important</strong>: In Crawl4AI, all strategy definitions should go inside the <code>CrawlerRunConfig</code>, not directly as a param in <code>arun()</code>. Here’s a full example:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> typing <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">List</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Product</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    name: <span class=\"hljs-built_in\">str</span>\n    price: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># 1. Define the LLM extraction strategy</span>\n    llm_strategy = LLMExtractionStrategy(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>, api_token=os.getenv(<span class=\"hljs-string\">'OPENAI_API_KEY'</span>)),\n        schema=Product.model_json_schema(), <span class=\"hljs-comment\"># Or use model_json_schema()</span>\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract all product objects with 'name' and 'price' from the content.\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">1000</span>,\n        overlap_rate=<span class=\"hljs-number\">0.0</span>,\n        apply_chunking=<span class=\"hljs-literal\">True</span>,\n        input_format=<span class=\"hljs-string\">\"markdown\"</span>,   <span class=\"hljs-comment\"># or \"html\", \"fit_markdown\"</span>\n        extra_args={<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0.0</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">800</span>}\n    )\n\n    <span class=\"hljs-comment\"># 2. Build the crawler config</span>\n    crawl_config = CrawlerRunConfig(\n        extraction_strategy=llm_strategy,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-comment\"># 3. Create a browser config if needed</span>\n    browser_cfg = BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=browser_cfg) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4. Let's say we want to crawl a single page</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n            config=crawl_config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-comment\"># 5. The extracted content is presumably JSON</span>\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Extracted items:\"</span>, data)\n\n            <span class=\"hljs-comment\"># 6. Show usage stats</span>\n            llm_strategy.show_usage()  <span class=\"hljs-comment\"># prints token usage</span>\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Error:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"6-chunking-details\">6. Chunking Details</h2>\n<h3 id=\"61-chunk_token_threshold\">6.1 <code>chunk_token_threshold</code></h3>\n<p>If your page is large, you might exceed your LLM’s context window. <strong><code>chunk_token_threshold</code></strong> sets the approximate max tokens per chunk. The library calculates word→token ratio using <code>word_token_rate</code> (often ~0.75 by default). If chunking is enabled (<code>apply_chunking=True</code>), the text is split into segments.</p>\n<h3 id=\"62-overlap_rate\">6.2 <code>overlap_rate</code></h3>\n<p>To keep context continuous across chunks, we can overlap them. E.g., <code>overlap_rate=0.1</code> means each subsequent chunk includes 10% of the previous chunk’s text. This is helpful if your needed info might straddle chunk boundaries.</p>\n<h3 id=\"63-performance-parallelism\">6.3 Performance &amp; Parallelism</h3>\n<p>By chunking, you can potentially process multiple chunks in parallel (depending on your concurrency settings and the LLM provider). This reduces total time if the site is huge or has many sections.</p>\n<hr>\n<h2 id=\"7-input-format\">7. Input Format</h2>\n<p>By default, <strong>LLMExtractionStrategy</strong> uses <code>input_format=\"markdown\"</code>, meaning the <strong>crawler’s final markdown</strong> is fed to the LLM. You can change to:</p>\n<ul>\n<li><strong><code>html</code></strong>: The cleaned HTML or raw HTML (depending on your crawler config) goes into the LLM.  </li>\n<li><strong><code>fit_markdown</code></strong>: If you used, for instance, <code>PruningContentFilter</code>, the “fit” version of the markdown is used. This can drastically reduce tokens if you trust the filter.  </li>\n<li><strong><code>markdown</code></strong>: Standard markdown output from the crawler’s <code>markdown_generator</code>.</li>\n</ul>\n<p>This setting is crucial: if the LLM instructions rely on HTML tags, pick <code>\"html\"</code>. If you prefer a text-based approach, pick <code>\"markdown\"</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">LLMExtractionStrategy(\n    <span class=\"hljs-comment\"># ...</span>\n    input_format=<span class=\"hljs-string\">\"html\"</span>,  <span class=\"hljs-comment\"># Instead of \"markdown\" or \"fit_markdown\"</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<hr>\n<h2 id=\"8-token-usage-show-usage\">8. Token Usage &amp; Show Usage</h2>\n<p>To keep track of tokens and cost, each chunk is processed with an LLM call. We record usage in:</p>\n<ul>\n<li><strong><code>usages</code></strong> (list): token usage per chunk or call.  </li>\n<li><strong><code>total_usage</code></strong>: sum of all chunk calls.  </li>\n<li><strong><code>show_usage()</code></strong>: prints a usage report (if the provider returns usage data).</li>\n</ul>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">llm_strategy = LLMExtractionStrategy(...)\n<span class=\"hljs-comment\"># ...</span>\nllm_strategy.show_usage()\n<span class=\"hljs-comment\"># e.g. “Total usage: 1241 tokens across 2 chunk calls”</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>If your model provider doesn't return usage info, these fields might be partial or empty.</p>\n<blockquote>\n<p><strong>Tip:</strong> <code>JsonCssExtractionStrategy.generate_schema()</code> also supports token usage tracking via an optional <code>usage</code> parameter. See <a href=\"../no-llm-strategies/#token-usage-tracking\">Token Usage Tracking in Schema Generation</a> for details.</p>\n</blockquote>\n<hr>\n<h2 id=\"9-example-building-a-knowledge-graph\">9. Example: Building a Knowledge Graph</h2>\n<p>Below is a snippet combining <strong><code>LLMExtractionStrategy</code></strong> with a Pydantic schema for a knowledge graph. Notice how we pass an <strong><code>instruction</code></strong> telling the model what to parse.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> os\n<span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> typing <span class=\"hljs-keyword\">import</span> <span class=\"hljs-type\">List</span>\n<span class=\"hljs-keyword\">from</span> pydantic <span class=\"hljs-keyword\">import</span> BaseModel, Field\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMExtractionStrategy\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Entity</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    name: <span class=\"hljs-built_in\">str</span>\n    description: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">Relationship</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    entity1: Entity\n    entity2: Entity\n    description: <span class=\"hljs-built_in\">str</span>\n    relation_type: <span class=\"hljs-built_in\">str</span>\n\n<span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">KnowledgeGraph</span>(<span class=\"hljs-title class_ inherited__\">BaseModel</span>):\n    entities: <span class=\"hljs-type\">List</span>[Entity]\n    relationships: <span class=\"hljs-type\">List</span>[Relationship]\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">main</span>():\n    <span class=\"hljs-comment\"># LLM extraction strategy</span>\n    llm_strat = LLMExtractionStrategy(\n        llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4\"</span>, api_token=os.getenv(<span class=\"hljs-string\">'OPENAI_API_KEY'</span>)),\n        schema=KnowledgeGraph.model_json_schema(),\n        extraction_type=<span class=\"hljs-string\">\"schema\"</span>,\n        instruction=<span class=\"hljs-string\">\"Extract entities and relationships from the content. Return valid JSON.\"</span>,\n        chunk_token_threshold=<span class=\"hljs-number\">1400</span>,\n        apply_chunking=<span class=\"hljs-literal\">True</span>,\n        input_format=<span class=\"hljs-string\">\"html\"</span>,\n        extra_args={<span class=\"hljs-string\">\"temperature\"</span>: <span class=\"hljs-number\">0.1</span>, <span class=\"hljs-string\">\"max_tokens\"</span>: <span class=\"hljs-number\">1500</span>}\n    )\n\n    crawl_config = CrawlerRunConfig(\n        extraction_strategy=llm_strat,\n        cache_mode=CacheMode.BYPASS\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(config=BrowserConfig(headless=<span class=\"hljs-literal\">True</span>)) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Example page</span>\n        url = <span class=\"hljs-string\">\"https://www.nbcnews.com/business\"</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(url=url, config=crawl_config)\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"--- LLM RAW RESPONSE ---\"</span>)\n        <span class=\"hljs-built_in\">print</span>(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"--- END LLM RAW RESPONSE ---\"</span>)\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            <span class=\"hljs-keyword\">with</span> <span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"kb_result.json\"</span>, <span class=\"hljs-string\">\"w\"</span>, encoding=<span class=\"hljs-string\">\"utf-8\"</span>) <span class=\"hljs-keyword\">as</span> f:\n                f.write(result.extracted_content)\n            llm_strat.show_usage()\n        <span class=\"hljs-keyword\">else</span>:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n\n<span class=\"hljs-keyword\">if</span> __name__ == <span class=\"hljs-string\">\"__main__\"</span>:\n    asyncio.run(main())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Observations</strong>:</p>\n<ul>\n<li><strong><code>extraction_type=\"schema\"</code></strong> ensures we get JSON fitting our <code>KnowledgeGraph</code>.  </li>\n<li><strong><code>input_format=\"html\"</code></strong> means we feed HTML to the model.  </li>\n<li><strong><code>instruction</code></strong> guides the model to output a structured knowledge graph.  </li>\n</ul>\n<hr>\n<h2 id=\"10-best-practices-caveats\">10. Best Practices &amp; Caveats</h2>\n<p>1. <strong>Cost &amp; Latency</strong>: LLM calls can be slow or expensive. Consider chunking or smaller coverage if you only need partial data.<br>\n2. <strong>Model Token Limits</strong>: If your page + instruction exceed the context window, chunking is essential.<br>\n3. <strong>Instruction Engineering</strong>: Well-crafted instructions can drastically improve output reliability.<br>\n4. <strong>Schema Strictness</strong>: <code>\"schema\"</code> extraction tries to parse the model output as JSON. If the model returns invalid JSON, partial extraction might happen, or you might get an error.<br>\n5. <strong>Parallel vs. Serial</strong>: The library can process multiple chunks in parallel, but you must watch out for rate limits on certain providers.<br>\n6. <strong>Check Output</strong>: Sometimes, an LLM might omit fields or produce extraneous text. You may want to post-validate with Pydantic or do additional cleanup.</p>\n<hr>\n<h2 id=\"11-conclusion\">11. Conclusion</h2>\n<p><strong>LLM-based extraction</strong> in Crawl4AI is <strong>provider-agnostic</strong>, letting you choose from hundreds of models via LiteLLM. It’s perfect for <strong>semantically complex</strong> tasks or generating advanced structures like knowledge graphs. However, it’s <strong>slower</strong> and potentially costlier than schema-based approaches. Keep these tips in mind:</p>\n<ul>\n<li>Put your LLM strategy <strong>in <code>CrawlerRunConfig</code></strong>.  </li>\n<li>Use <strong><code>input_format</code></strong> to pick which form (markdown, HTML, fit_markdown) the LLM sees.  </li>\n<li>Tweak <strong><code>chunk_token_threshold</code></strong>, <strong><code>overlap_rate</code></strong>, and <strong><code>apply_chunking</code></strong> to handle large content efficiently.  </li>\n<li>Monitor token usage with <code>show_usage()</code>.</li>\n</ul>\n<p>If your site’s data is consistent or repetitive, consider <a href=\"../no-llm-strategies/\"><code>JsonCssExtractionStrategy</code></a> first for speed and simplicity. But if you need an <strong>AI-driven</strong> approach, <code>LLMExtractionStrategy</code> offers a flexible, multi-provider solution for extracting structured JSON from any website.</p>\n<p><strong>Next Steps</strong>:</p>\n<p>1. <strong>Experiment with Different Providers</strong><br>\n   - Try switching the <code>provider</code> (e.g., <code>\"ollama/llama2\"</code>, <code>\"openai/gpt-4o\"</code>, etc.) to see differences in speed, accuracy, or cost.<br>\n   - Pass different <code>extra_args</code> like <code>temperature</code>, <code>top_p</code>, and <code>max_tokens</code> to fine-tune your results.</p>\n<p>2. <strong>Performance Tuning</strong><br>\n   - If pages are large, tweak <code>chunk_token_threshold</code>, <code>overlap_rate</code>, or <code>apply_chunking</code> to optimize throughput.<br>\n   - Check the usage logs with <code>show_usage()</code> to keep an eye on token consumption and identify potential bottlenecks.</p>\n<p>3. <strong>Validate Outputs</strong><br>\n   - If using <code>extraction_type=\"schema\"</code>, parse the LLM’s JSON with a Pydantic model for a final validation step.<br>\n   - Log or handle any parse errors gracefully, especially if the model occasionally returns malformed JSON.</p>\n<p>4. <strong>Explore Hooks &amp; Automation</strong><br>\n   - Integrate LLM extraction with <a href=\"../../advanced/hooks-auth/\">hooks</a> for complex pre/post-processing.<br>\n   - Use a multi-step pipeline: crawl, filter, LLM-extract, then store or index results for further analysis.</p>\n<p><strong>Last Updated</strong>: 2025-01-01</p>\n<hr>\n<p>That’s it for <strong>Extracting JSON (LLM)</strong>—now you can harness AI to parse, classify, or reorganize data on the web. Happy crawling!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/extraction/no-llm-strategies/",
    "title": "LLM-Free Strategies - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"extracting-json-no-llm\">Extracting JSON (No LLM)</h1>\n<p>One of Crawl4AI's <strong>most powerful</strong> features is extracting <strong>structured JSON</strong> from websites <strong>without</strong> relying on large language models. Crawl4AI offers several strategies for LLM-free extraction:</p>\n<ol>\n<li><strong>Schema-based extraction</strong> with CSS or XPath selectors via <code>JsonCssExtractionStrategy</code> and <code>JsonXPathExtractionStrategy</code></li>\n<li><strong>Regular expression extraction</strong> with <code>RegexExtractionStrategy</code> for fast pattern matching</li>\n</ol>\n<p>These approaches let you extract data instantly—even from complex or nested HTML structures—without the cost, latency, or environmental impact of an LLM.</p>\n<p><strong>Why avoid LLM for basic extractions?</strong></p>\n<ol>\n<li><strong>Faster &amp; Cheaper</strong>: No API calls or GPU overhead.  </li>\n<li><strong>Lower Carbon Footprint</strong>: LLM inference can be energy-intensive. Pattern-based extraction is practically carbon-free.  </li>\n<li><strong>Precise &amp; Repeatable</strong>: CSS/XPath selectors and regex patterns do exactly what you specify. LLM outputs can vary or hallucinate.  </li>\n<li><strong>Scales Readily</strong>: For thousands of pages, pattern-based extraction runs quickly and in parallel.</li>\n</ol>\n<p>Below, we'll explore how to craft these schemas and use them with <strong>JsonCssExtractionStrategy</strong> (or <strong>JsonXPathExtractionStrategy</strong> if you prefer XPath). We'll also highlight advanced features like <strong>nested fields</strong> and <strong>base element attributes</strong>.</p>\n<hr>\n<h2 id=\"1-intro-to-schema-based-extraction\">1. Intro to Schema-Based Extraction</h2>\n<p>A schema defines:</p>\n<ol>\n<li>A <strong>base selector</strong> that identifies each \"container\" element on the page (e.g., a product row, a blog post card).  </li>\n<li><strong>Fields</strong> describing which CSS/XPath selectors to use for each piece of data you want to capture (text, attribute, HTML block, etc.).  </li>\n<li><strong>Nested</strong> or <strong>list</strong> types for repeated or hierarchical structures.  </li>\n</ol>\n<p>For example, if you have a list of products, each one might have a name, price, reviews, and \"related products.\" This approach is faster and more reliable than an LLM for consistent, structured pages.</p>\n<hr>\n<h2 id=\"2-simple-example-crypto-prices\">2. Simple Example: Crypto Prices</h2>\n<p>Let's begin with a <strong>simple</strong> schema-based extraction using the <code>JsonCssExtractionStrategy</code>. Below is a snippet that extracts cryptocurrency prices from a site (similar to the legacy Coinbase example). Notice we <strong>don't</strong> call any LLM:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, CacheMode\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_crypto_prices</span>():\n    <span class=\"hljs-comment\"># 1. Define a simple extraction schema</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Crypto Prices\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"div.crypto-row\"</span>,    <span class=\"hljs-comment\"># Repeated elements</span>\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"coin_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"h2.coin-name\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"price\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\"span.coin-price\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            }\n        ]\n    }\n\n    <span class=\"hljs-comment\"># 2. Create the extraction strategy</span>\n    extraction_strategy = JsonCssExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n    <span class=\"hljs-comment\"># 3. Set up your crawler config (if needed)</span>\n    config = CrawlerRunConfig(\n        <span class=\"hljs-comment\"># e.g., pass js_code or wait_for if the page is dynamic</span>\n        <span class=\"hljs-comment\"># wait_for=\"css:.crypto-row:nth-child(20)\"</span>\n        cache_mode = CacheMode.BYPASS,\n        extraction_strategy=extraction_strategy,\n    )\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># 4. Run the crawl and extraction</span>\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/crypto-prices\"</span>,\n\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># 5. Parse the extracted JSON</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span> coin entries\"</span>)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(data[<span class=\"hljs-number\">0</span>], indent=<span class=\"hljs-number\">2</span>) <span class=\"hljs-keyword\">if</span> data <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">\"No data found\"</span>)\n\nasyncio.run(extract_crypto_prices())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Highlights</strong>:</p>\n<ul>\n<li><strong><code>baseSelector</code></strong>: Tells us where each \"item\" (crypto row) is.</li>\n<li><strong><code>fields</code></strong>: Two fields (<code>coin_name</code>, <code>price</code>) using simple CSS selectors.</li>\n<li>Each field defines a <strong><code>type</code></strong> (e.g., <code>text</code>, <code>attribute</code>, <code>html</code>, <code>regex</code>, etc.).</li>\n<li>Optional keys: <strong><code>transform</code></strong>, <strong><code>default</code></strong>, <strong><code>attribute</code></strong>, <strong><code>pattern</code></strong>, and <strong><code>source</code></strong> (for sibling data — see <a href=\"#sibling-data\">Extracting Sibling Data</a>).</li>\n</ul>\n<p>No LLM is needed, and the performance is <strong>near-instant</strong> for hundreds or thousands of items.</p>\n<hr>\n<h3 id=\"xpath-example-with-raw-html\"><strong>XPath Example with <code>raw://</code> HTML</strong></h3>\n<p>Below is a short example demonstrating <strong>XPath</strong> extraction plus the <strong><code>raw://</code></strong> scheme. We'll pass a <strong>dummy HTML</strong> directly (no network request) and define the extraction strategy in <code>CrawlerRunConfig</code>.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonXPathExtractionStrategy\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_crypto_prices_xpath</span>():\n    <span class=\"hljs-comment\"># 1. Minimal dummy HTML with some repeating rows</span>\n    dummy_html = <span class=\"hljs-string\">\"\"\"\n    &lt;html&gt;\n      &lt;body&gt;\n        &lt;div class='crypto-row'&gt;\n          &lt;h2 class='coin-name'&gt;Bitcoin&lt;/h2&gt;\n          &lt;span class='coin-price'&gt;$28,000&lt;/span&gt;\n        &lt;/div&gt;\n        &lt;div class='crypto-row'&gt;\n          &lt;h2 class='coin-name'&gt;Ethereum&lt;/h2&gt;\n          &lt;span class='coin-price'&gt;$1,800&lt;/span&gt;\n        &lt;/div&gt;\n      &lt;/body&gt;\n    &lt;/html&gt;\n    \"\"\"</span>\n\n    <span class=\"hljs-comment\"># 2. Define the JSON schema (XPath version)</span>\n    schema = {\n        <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"Crypto Prices via XPath\"</span>,\n        <span class=\"hljs-string\">\"baseSelector\"</span>: <span class=\"hljs-string\">\"//div[@class='crypto-row']\"</span>,\n        <span class=\"hljs-string\">\"fields\"</span>: [\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"coin_name\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".//h2[@class='coin-name']\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            },\n            {\n                <span class=\"hljs-string\">\"name\"</span>: <span class=\"hljs-string\">\"price\"</span>,\n                <span class=\"hljs-string\">\"selector\"</span>: <span class=\"hljs-string\">\".//span[@class='coin-price']\"</span>,\n                <span class=\"hljs-string\">\"type\"</span>: <span class=\"hljs-string\">\"text\"</span>\n            }\n        ]\n    }\n\n    <span class=\"hljs-comment\"># 3. Place the strategy in the CrawlerRunConfig</span>\n    config = CrawlerRunConfig(\n        extraction_strategy=JsonXPathExtractionStrategy(schema, verbose=<span class=\"hljs-literal\">True</span>)\n    )\n\n    <span class=\"hljs-comment\"># 4. Use raw:// scheme to pass dummy_html directly</span>\n    raw_url = <span class=\"hljs-string\">f\"raw://<span class=\"hljs-subst\">{dummy_html}</span>\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=raw_url,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span> coin rows\"</span>)\n        <span class=\"hljs-keyword\">if</span> data:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"First item:\"</span>, data[<span class=\"hljs-number\">0</span>])\n\nasyncio.run(extract_crypto_prices_xpath())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p><strong>Key Points</strong>:</p>\n<ol>\n<li><strong><code>JsonXPathExtractionStrategy</code></strong> is used instead of <code>JsonCssExtractionStrategy</code>.  </li>\n<li><strong><code>baseSelector</code></strong> and each field's <code>\"selector\"</code> use <strong>XPath</strong> instead of CSS.  </li>\n<li><strong><code>raw://</code></strong> lets us pass <code>dummy_html</code> with no real network request—handy for local testing.  </li>\n<li>Everything (including the extraction strategy) is in <strong><code>CrawlerRunConfig</code></strong>.  </li>\n</ol>\n<p>That's how you keep the config self-contained, illustrate <strong>XPath</strong> usage, and demonstrate the <strong>raw</strong> scheme for direct HTML input—all while avoiding the old approach of passing <code>extraction_strategy</code> directly to <code>arun()</code>.</p>\n<hr>\n<h2 id=\"3-advanced-schema-nested-structures\">3. Advanced Schema &amp; Nested Structures</h2>\n<p>Real sites often have <strong>nested</strong> or repeated data—like categories containing products, which themselves have a list of reviews or features. For that, we can define <strong>nested</strong> or <strong>list</strong> (and even <strong>nested_list</strong>) fields.</p>\n<h3 id=\"sample-e-commerce-html\">Sample E-Commerce HTML</h3>\n<p>We have a <strong>sample e-commerce</strong> HTML file on GitHub (example):\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/examples/sample_ecommerce.html\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\nThis snippet includes categories, products, features, reviews, and related items. Let's see how to define a schema that fully captures that structure <strong>without LLM</strong>.<p></p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"E-commerce Product Catalog\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.category\"</span>,\n    <span class=\"hljs-comment\"># (1) We can define optional baseFields if we want to extract attributes </span>\n    <span class=\"hljs-comment\"># from the category container</span>\n    <span class=\"hljs-string\">\"baseFields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"data_cat_id\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"data-cat-id\"</span><span class=\"hljs-punctuation\">}</span>, \n    <span class=\"hljs-punctuation\">]</span>,\n    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span>\n            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"category_name\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h2.category-name\"</span>,\n            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n        <span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span>\n            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"products\"</span>,\n            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.product\"</span>,\n            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested_list\"</span>,    <span class=\"hljs-comment\"># repeated sub-objects</span>\n            <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"name\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h3.product-name\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"price\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.product-price\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"details\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.product-details\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested\"</span>,  <span class=\"hljs-comment\"># single sub-object</span>\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"brand\"</span>,\n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.brand\"</span>,\n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"model\"</span>,\n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.model\"</span>,\n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"features\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"ul.product-features li\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"feature\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span> \n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"reviews\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"div.review\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"nested_list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"reviewer\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.reviewer\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"rating\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.rating\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"comment\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.review-text\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>,\n                <span class=\"hljs-punctuation\">{</span>\n                    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"related_products\"</span>,\n                    <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"ul.related-products li\"</span>,\n                    <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"list\"</span>,\n                    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"name\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.related-name\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>,\n                        <span class=\"hljs-punctuation\">{</span>\n                            <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"price\"</span>, \n                            <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.related-price\"</span>, \n                            <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>\n                        <span class=\"hljs-punctuation\">}</span>\n                    <span class=\"hljs-punctuation\">]</span>\n                <span class=\"hljs-punctuation\">}</span>\n            <span class=\"hljs-punctuation\">]</span>\n        <span class=\"hljs-punctuation\">}</span>\n    <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Key Takeaways:</p>\n<ul>\n<li><strong>Nested vs. List</strong>:  </li>\n<li><strong><code>type: \"nested\"</code></strong> means a <strong>single</strong> sub-object (like <code>details</code>).  </li>\n<li><strong><code>type: \"list\"</code></strong> means multiple items that are <strong>simple</strong> dictionaries or single text fields.  </li>\n<li><strong><code>type: \"nested_list\"</code></strong> means repeated <strong>complex</strong> objects (like <code>products</code> or <code>reviews</code>).</li>\n<li><strong>Base Fields</strong>: We can extract <strong>attributes</strong> from the container element via <code>\"baseFields\"</code>. For instance, <code>\"data_cat_id\"</code> might be <code>data-cat-id=\"elect123\"</code>.  </li>\n<li><strong>Transforms</strong>: We can also define a <code>transform</code> if we want to lower/upper case, strip whitespace, or even run a custom function.</li>\n</ul>\n<h3 id=\"running-the-extraction\">Running the Extraction</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n\necommerce_schema = {\n    <span class=\"hljs-comment\"># ... the advanced schema from above ...</span>\n}\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_ecommerce_data</span>():\n    strategy = JsonCssExtractionStrategy(ecommerce_schema, verbose=<span class=\"hljs-literal\">True</span>)\n\n    config = CrawlerRunConfig()\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler(verbose=<span class=\"hljs-literal\">True</span>) <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/examples/sample_ecommerce.html\"</span>,\n            extraction_strategy=strategy,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> <span class=\"hljs-keyword\">not</span> result.success:\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Crawl failed:\"</span>, result.error_message)\n            <span class=\"hljs-keyword\">return</span>\n\n        <span class=\"hljs-comment\"># Parse the JSON output</span>\n        data = json.loads(result.extracted_content)\n        <span class=\"hljs-built_in\">print</span>(json.dumps(data, indent=<span class=\"hljs-number\">2</span>) <span class=\"hljs-keyword\">if</span> data <span class=\"hljs-keyword\">else</span> <span class=\"hljs-string\">\"No data found.\"</span>)\n\nasyncio.run(extract_ecommerce_data())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>If all goes well, you get a <strong>structured</strong> JSON array with each \"category,\" containing an array of <code>products</code>. Each product includes <code>details</code>, <code>features</code>, <code>reviews</code>, etc. All of that <strong>without</strong> an LLM.</p>\n<hr>\n<h2 id=\"4-regexextractionstrategy-fast-pattern-based-extraction\">4. RegexExtractionStrategy - Fast Pattern-Based Extraction</h2>\n<p>Crawl4AI now offers a powerful new zero-LLM extraction strategy: <code>RegexExtractionStrategy</code>. This strategy provides lightning-fast extraction of common data types like emails, phone numbers, URLs, dates, and more using pre-compiled regular expressions.</p>\n<h3 id=\"key-features\">Key Features</h3>\n<ul>\n<li><strong>Zero LLM Dependency</strong>: Extracts data without any AI model calls</li>\n<li><strong>Blazing Fast</strong>: Uses pre-compiled regex patterns for maximum performance</li>\n<li><strong>Built-in Patterns</strong>: Includes ready-to-use patterns for common data types</li>\n<li><strong>Custom Patterns</strong>: Add your own regex patterns for domain-specific extraction</li>\n<li><strong>LLM-Assisted Pattern Generation</strong>: Optionally use an LLM once to generate optimized patterns, then reuse them without further LLM calls</li>\n</ul>\n<h3 id=\"simple-example-extracting-common-entities\">Simple Example: Extracting Common Entities</h3>\n<p>The easiest way to start is by using the built-in pattern catalog:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_with_regex</span>():\n    <span class=\"hljs-comment\"># Create a strategy using built-in patterns for URLs and currencies</span>\n    strategy = RegexExtractionStrategy(\n        pattern = RegexExtractionStrategy.Url | RegexExtractionStrategy.Currency\n    )\n\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data[:<span class=\"hljs-number\">5</span>]:  <span class=\"hljs-comment\"># Show first 5 matches</span>\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"<span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'label'</span>]}</span>: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total matches: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span>\"</span>)\n\nasyncio.run(extract_with_regex())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"available-built-in-patterns\">Available Built-in Patterns</h3>\n<p><code>RegexExtractionStrategy</code> provides these common patterns as IntFlag attributes for easy combining:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Use individual patterns</span>\nstrategy = RegexExtractionStrategy(pattern=RegexExtractionStrategy.Email)\n\n<span class=\"hljs-comment\"># Combine multiple patterns</span>\nstrategy = RegexExtractionStrategy(\n    pattern = (\n        RegexExtractionStrategy.Email | \n        RegexExtractionStrategy.PhoneUS | \n        RegexExtractionStrategy.Url\n    )\n)\n\n<span class=\"hljs-comment\"># Use all available patterns</span>\nstrategy = RegexExtractionStrategy(pattern=RegexExtractionStrategy.All)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Available patterns include:\n- <code>Email</code> - Email addresses\n- <code>PhoneIntl</code> - International phone numbers\n- <code>PhoneUS</code> - US-format phone numbers\n- <code>Url</code> - HTTP/HTTPS URLs\n- <code>IPv4</code> - IPv4 addresses\n- <code>IPv6</code> - IPv6 addresses\n- <code>Uuid</code> - UUIDs\n- <code>Currency</code> - Currency values (USD, EUR, etc.)\n- <code>Percentage</code> - Percentage values\n- <code>Number</code> - Numeric values\n- <code>DateIso</code> - ISO format dates\n- <code>DateUS</code> - US format dates\n- <code>Time24h</code> - 24-hour format times\n- <code>PostalUS</code> - US postal codes\n- <code>PostalUK</code> - UK postal codes\n- <code>HexColor</code> - HTML hex color codes\n- <code>TwitterHandle</code> - Twitter handles\n- <code>Hashtag</code> - Hashtags\n- <code>MacAddr</code> - MAC addresses\n- <code>Iban</code> - International bank account numbers\n- <code>CreditCard</code> - Credit card numbers</p>\n<h3 id=\"custom-pattern-example\">Custom Pattern Example</h3>\n<p>For more targeted extraction, you can provide custom patterns:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_prices</span>():\n    <span class=\"hljs-comment\"># Define a custom pattern for US Dollar prices</span>\n    price_pattern = {<span class=\"hljs-string\">\"usd_price\"</span>: <span class=\"hljs-string\">r\"\\$\\s?\\d{1,3}(?:,\\d{3})*(?:\\.\\d{2})?\"</span>}\n\n    <span class=\"hljs-comment\"># Create strategy with custom pattern</span>\n    strategy = RegexExtractionStrategy(custom=price_pattern)\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://www.example.com/products\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Found price: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n\nasyncio.run(extract_prices())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"llm-assisted-pattern-generation\">LLM-Assisted Pattern Generation</h3>\n<p>For complex or site-specific patterns, you can use an LLM once to generate an optimized pattern, then save and reuse it without further LLM calls:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> json\n<span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> pathlib <span class=\"hljs-keyword\">import</span> Path\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    AsyncWebCrawler,\n    CrawlerRunConfig,\n    RegexExtractionStrategy,\n    LLMConfig\n)\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_with_generated_pattern</span>():\n    cache_dir = Path(<span class=\"hljs-string\">\"./pattern_cache\"</span>)\n    cache_dir.mkdir(exist_ok=<span class=\"hljs-literal\">True</span>)\n    pattern_file = cache_dir / <span class=\"hljs-string\">\"price_pattern.json\"</span>\n\n    <span class=\"hljs-comment\"># 1. Generate or load pattern</span>\n    <span class=\"hljs-keyword\">if</span> pattern_file.exists():\n        pattern = json.load(pattern_file.<span class=\"hljs-built_in\">open</span>())\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Using cached pattern: <span class=\"hljs-subst\">{pattern}</span>\"</span>)\n    <span class=\"hljs-keyword\">else</span>:\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"Generating pattern via LLM...\"</span>)\n\n        <span class=\"hljs-comment\"># Configure LLM</span>\n        llm_config = LLMConfig(\n            provider=<span class=\"hljs-string\">\"openai/gpt-4o-mini\"</span>,\n            api_token=<span class=\"hljs-string\">\"env:OPENAI_API_KEY\"</span>,\n        )\n\n        <span class=\"hljs-comment\"># Get sample HTML for context</span>\n        <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n            result = <span class=\"hljs-keyword\">await</span> crawler.arun(<span class=\"hljs-string\">\"https://example.com/products\"</span>)\n            html = result.markdown.fit_html\n\n        <span class=\"hljs-comment\"># Generate pattern (one-time LLM usage)</span>\n        pattern = RegexExtractionStrategy.generate_pattern(\n            label=<span class=\"hljs-string\">\"price\"</span>,\n            html=html,\n            query=<span class=\"hljs-string\">\"Product prices in USD format\"</span>,\n            llm_config=llm_config,\n        )\n\n        <span class=\"hljs-comment\"># Cache pattern for future use</span>\n        json.dump(pattern, pattern_file.<span class=\"hljs-built_in\">open</span>(<span class=\"hljs-string\">\"w\"</span>), indent=<span class=\"hljs-number\">2</span>)\n\n    <span class=\"hljs-comment\"># 2. Use pattern for extraction (no LLM calls)</span>\n    strategy = RegexExtractionStrategy(custom=pattern)\n    config = CrawlerRunConfig(extraction_strategy=strategy)\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        result = <span class=\"hljs-keyword\">await</span> crawler.arun(\n            url=<span class=\"hljs-string\">\"https://example.com/products\"</span>,\n            config=config\n        )\n\n        <span class=\"hljs-keyword\">if</span> result.success:\n            data = json.loads(result.extracted_content)\n            <span class=\"hljs-keyword\">for</span> item <span class=\"hljs-keyword\">in</span> data[:<span class=\"hljs-number\">10</span>]:\n                <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Extracted: <span class=\"hljs-subst\">{item[<span class=\"hljs-string\">'value'</span>]}</span>\"</span>)\n            <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total matches: <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(data)}</span>\"</span>)\n\nasyncio.run(extract_with_generated_pattern())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>This pattern allows you to:\n1. Use an LLM once to generate a highly optimized regex for your specific site\n2. Save the pattern to disk for reuse \n3. Extract data using only regex (no further LLM calls) in production</p>\n<h3 id=\"extraction-results-format\">Extraction Results Format</h3>\n<p>The <code>RegexExtractionStrategy</code> returns results in a consistent format:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-perl\">[\n  {\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-string\">\"label\"</span>: <span class=\"hljs-string\">\"email\"</span>,\n    <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"contact@example.com\"</span>,\n    <span class=\"hljs-string\">\"span\"</span>: [<span class=\"hljs-number\">145</span>, <span class=\"hljs-number\">163</span>]\n  },\n  {\n    <span class=\"hljs-string\">\"url\"</span>: <span class=\"hljs-string\">\"https://example.com\"</span>,\n    <span class=\"hljs-string\">\"label\"</span>: <span class=\"hljs-string\">\"url\"</span>,\n    <span class=\"hljs-string\">\"value\"</span>: <span class=\"hljs-string\">\"https://support.example.com\"</span>,\n    <span class=\"hljs-string\">\"span\"</span>: [<span class=\"hljs-number\">210</span>, <span class=\"hljs-number\">235</span>]\n  }\n]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Each match includes:\n- <code>url</code>: The source URL\n- <code>label</code>: The pattern name that matched (e.g., \"email\", \"phone_us\")\n- <code>value</code>: The extracted text\n- <code>span</code>: The start and end positions in the source content</p>\n<hr>\n<h2 id=\"5-why-no-llm-is-often-better\">5. Why \"No LLM\" Is Often Better</h2>\n<ol>\n<li><strong>Zero Hallucination</strong>: Pattern-based extraction doesn't guess text. It either finds it or not.  </li>\n<li><strong>Guaranteed Structure</strong>: The same schema or regex yields consistent JSON across many pages, so your downstream pipeline can rely on stable keys.  </li>\n<li><strong>Speed</strong>: LLM-based extraction can be 10–1000x slower for large-scale crawling.  </li>\n<li><strong>Scalable</strong>: Adding or updating a field is a matter of adjusting the schema or regex, not re-tuning a model.</li>\n</ol>\n<p><strong>When might you consider an LLM?</strong> Possibly if the site is extremely unstructured or you want AI summarization. But always try a schema or regex approach first for repeated or consistent data patterns.</p>\n<hr>\n<h2 id=\"6-base-element-attributes-additional-fields\">6. Base Element Attributes &amp; Additional Fields</h2>\n<p>It's easy to <strong>extract attributes</strong> (like <code>href</code>, <code>src</code>, or <code>data-xxx</code>) from your base or nested elements using:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-attr\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">,</span>\n  <span class=\"hljs-attr\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-literal\"><span class=\"hljs-keyword\">null</span></span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>You can define them in <strong><code>baseFields</code></strong> (extracted from the main container element) or in each field's sub-lists. This is especially helpful if you need an item's link or ID stored in the parent <code>&lt;div&gt;</code>.</p>\n<hr>\n<h2 id=\"7-putting-it-all-together-larger-example\">7. Putting It All Together: Larger Example</h2>\n<p>Consider a blog site. We have a schema that extracts the <strong>URL</strong> from each post card (via <code>baseFields</code> with an <code>\"attribute\": \"href\"</code>), plus the title, date, summary, and author:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n  <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"Blog Posts\"</span>,\n  <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"a.blog-post-card\"</span>,\n  <span class=\"hljs-string\">\"baseFields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"post_url\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">]</span>,\n  <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"h2.post-title\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"No Title\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"date\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"time.post-date\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"summary\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"p.post-summary\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"author\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.post-author\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"default\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"\"</span><span class=\"hljs-punctuation\">}</span>\n  <span class=\"hljs-punctuation\">]</span>\n<span class=\"hljs-punctuation\">}</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Then run with <code>JsonCssExtractionStrategy(schema)</code> to get an array of blog post objects, each with <code>\"post_url\"</code>, <code>\"title\"</code>, <code>\"date\"</code>, <code>\"summary\"</code>, <code>\"author\"</code>.</p>\n<hr>\n<h2 id=\"sibling-data\">8. Extracting Sibling Data with <code>source</code></h2>\n<p>Some websites split a single logical item across <strong>sibling elements</strong> rather than nesting everything inside one container. A classic example is Hacker News, where each submission spans two adjacent <code>&lt;tr&gt;</code> rows:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\"><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"athing submission\"</span>&gt;</span>  <span class=\"hljs-comment\">&lt;!-- rank, title, url --&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">span</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"rank\"</span>&gt;</span>1.<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">span</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">span</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"titleline\"</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">a</span> <span class=\"hljs-attr\">href</span>=<span class=\"hljs-string\">\"https://example.com\"</span>&gt;</span>Example Title<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">a</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">span</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span>                             <span class=\"hljs-comment\">&lt;!-- score, author, comments (sibling!) --&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"subtext\"</span>&gt;</span>\n    <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">span</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"score\"</span>&gt;</span>100 points<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">span</span>&gt;</span>\n    <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">a</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"hnuser\"</span>&gt;</span>johndoe<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">a</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Normally, field selectors only search <strong>descendants</strong> of the base element — siblings are unreachable. The <code>source</code> field key solves this by navigating to a sibling element before running the selector.</p>\n<h3 id=\"syntax\">Syntax</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-json\"><span class=\"hljs-attr\">\"source\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"+ &lt;selector&gt;\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<ul>\n<li><strong><code>+ tr</code></strong> — next sibling <code>&lt;tr&gt;</code></li>\n<li><strong><code>+ div.details</code></strong> — next sibling <code>&lt;div&gt;</code> with class <code>details</code></li>\n<li><strong><code>+ .subtext</code></strong> — next sibling with class <code>subtext</code></li>\n</ul>\n<h3 id=\"example-hacker-news\">Example: Hacker News</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> <span class=\"hljs-punctuation\">{</span>\n    <span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"HN Submissions\"</span>,\n    <span class=\"hljs-string\">\"baseSelector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"tr.athing.submission\"</span>,\n    <span class=\"hljs-string\">\"fields\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-punctuation\">[</span>\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"rank\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.rank\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"title\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.titleline a\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span><span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"url\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.titleline a\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"attribute\"</span>, <span class=\"hljs-string\">\"attribute\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"href\"</span><span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"score\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"span.score\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"source\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"+ tr\"</span><span class=\"hljs-punctuation\">}</span>,\n        <span class=\"hljs-punctuation\">{</span><span class=\"hljs-string\">\"name\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"author\"</span>, <span class=\"hljs-string\">\"selector\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"a.hnuser\"</span>, <span class=\"hljs-string\">\"type\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"text\"</span>, <span class=\"hljs-string\">\"source\"</span><span class=\"hljs-punctuation\">:</span> <span class=\"hljs-string\">\"+ tr\"</span><span class=\"hljs-punctuation\">}</span>,\n    <span class=\"hljs-punctuation\">]</span>,\n<span class=\"hljs-punctuation\">}</span>\n\nstrategy <span class=\"hljs-punctuation\">=</span> JsonCssExtractionStrategy<span class=\"hljs-punctuation\">(</span><span class=\"hljs-keyword\">schema</span><span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The <code>score</code> and <code>author</code> fields first navigate to the next sibling <code>&lt;tr&gt;</code>, then run their selectors inside that element. Fields without <code>source</code> work as before — searching descendants of the base element.</p>\n<p><code>source</code> works with all field types (<code>text</code>, <code>attribute</code>, <code>nested</code>, <code>list</code>, etc.) and with both <code>JsonCssExtractionStrategy</code> and <code>JsonXPathExtractionStrategy</code>. If the sibling isn't found, the field returns its <code>default</code> value.</p>\n<hr>\n<h2 id=\"9-tips-best-practices\">9. Tips &amp; Best Practices</h2>\n<ol>\n<li><strong>Inspect the DOM</strong> in Chrome DevTools or Firefox's Inspector to find stable selectors.  </li>\n<li><strong>Start Simple</strong>: Verify you can extract a single field. Then add complexity like nested objects or lists.  </li>\n<li><strong>Test</strong> your schema on partial HTML or a test page before a big crawl.  </li>\n<li><strong>Combine with JS Execution</strong> if the site loads content dynamically. You can pass <code>js_code</code> or <code>wait_for</code> in <code>CrawlerRunConfig</code>.  </li>\n<li><strong>Look at Logs</strong> when <code>verbose=True</code>: if your selectors are off or your schema is malformed, it'll often show warnings.  </li>\n<li><strong>Use baseFields</strong> if you need attributes from the container element (e.g., <code>href</code>, <code>data-id</code>), especially for the \"parent\" item.  </li>\n<li><strong>Performance</strong>: For large pages, make sure your selectors are as narrow as possible.</li>\n<li><strong>Consider Using Regex First</strong>: For simple data types like emails, URLs, and dates, <code>RegexExtractionStrategy</code> is often the fastest approach.</li>\n</ol>\n<hr>\n<h2 id=\"10-schema-generation-utility\">10. Schema Generation Utility</h2>\n<p>While manually crafting schemas is powerful and precise, Crawl4AI now offers a convenient utility to <strong>automatically generate</strong> extraction schemas using LLM. This is particularly useful when:</p>\n<ol>\n<li>You're dealing with a new website structure and want a quick starting point</li>\n<li>You need to extract complex nested data structures</li>\n<li>You want to avoid the learning curve of CSS/XPath selector syntax</li>\n</ol>\n<h3 id=\"using-the-schema-generator\">Using the Schema Generator</h3>\n<p>The schema generator is available as a static method on both <code>JsonCssExtractionStrategy</code> and <code>JsonXPathExtractionStrategy</code>. You can choose between OpenAI's GPT-4 or the open-source Ollama for schema generation:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy, JsonXPathExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LLMConfig\n\n<span class=\"hljs-comment\"># Sample HTML with product information</span>\nhtml = <span class=\"hljs-string\">\"\"\"\n&lt;div class=\"product-card\"&gt;\n    &lt;h2 class=\"title\"&gt;Gaming Laptop&lt;/h2&gt;\n    &lt;div class=\"price\"&gt;$999.99&lt;/div&gt;\n    &lt;div class=\"specs\"&gt;\n        &lt;ul&gt;\n            &lt;li&gt;16GB RAM&lt;/li&gt;\n            &lt;li&gt;1TB SSD&lt;/li&gt;\n        &lt;/ul&gt;\n    &lt;/div&gt;\n&lt;/div&gt;\n\"\"\"</span>\n\n<span class=\"hljs-comment\"># Option 1: Using OpenAI (requires API token)</span>\ncss_schema = JsonCssExtractionStrategy.generate_schema(\n    html,\n    schema_type=<span class=\"hljs-string\">\"css\"</span>,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"openai/gpt-4o\"</span>,api_token=<span class=\"hljs-string\">\"your-openai-token\"</span>)\n)\n\n<span class=\"hljs-comment\"># Option 2: Using Ollama (open source, no token needed)</span>\nxpath_schema = JsonXPathExtractionStrategy.generate_schema(\n    html,\n    schema_type=<span class=\"hljs-string\">\"xpath\"</span>,\n    llm_config = LLMConfig(provider=<span class=\"hljs-string\">\"ollama/llama3.3\"</span>, api_token=<span class=\"hljs-literal\">None</span>)  <span class=\"hljs-comment\"># Not needed for Ollama</span>\n)\n\n<span class=\"hljs-comment\"># Use the generated schema for fast, repeated extractions</span>\nstrategy = JsonCssExtractionStrategy(css_schema)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"schema-validation\">Schema Validation</h3>\n<p>By default, <code>generate_schema</code> <strong>validates</strong> the generated schema against the HTML to ensure that it actually extracts the data you expect. If the schema doesn't produce results, it automatically refines the selectors before returning.</p>\n<p>You can control this with the <code>validate</code> parameter:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-graphql\"><span class=\"hljs-comment\"># Default: validated (recommended)</span>\n<span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> JsonCssExtractionStrategy.generate_schema<span class=\"hljs-punctuation\">(</span>\n    url<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Extract each story: title, url, score, author\"</span>,\n<span class=\"hljs-punctuation\">)</span>\n\n<span class=\"hljs-comment\"># Skip validation if you want raw LLM output</span>\n<span class=\"hljs-keyword\">schema</span> <span class=\"hljs-punctuation\">=</span> JsonCssExtractionStrategy.generate_schema<span class=\"hljs-punctuation\">(</span>\n    url<span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n    <span class=\"hljs-keyword\">query</span><span class=\"hljs-punctuation\">=</span><span class=\"hljs-string\">\"Extract each story: title, url, score, author\"</span>,\n    validate<span class=\"hljs-punctuation\">=</span><span class=\"hljs-literal\">False</span>,\n<span class=\"hljs-punctuation\">)</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The generator also understands sibling layouts — for sites like Hacker News where data is split across sibling elements, it will automatically use the <a href=\"#sibling-data\"><code>source</code> field</a> to reach sibling data.</p>\n<h3 id=\"token-usage-tracking\">Token Usage Tracking</h3>\n<p><code>generate_schema</code> may make multiple LLM calls internally (field inference, schema generation, validation retries). To track the total token consumption across all of these calls, pass a <code>TokenUsage</code> accumulator:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> JsonCssExtractionStrategy\n<span class=\"hljs-keyword\">from</span> crawl4ai.models <span class=\"hljs-keyword\">import</span> TokenUsage\n\nusage = TokenUsage()\n\nschema = JsonCssExtractionStrategy.generate_schema(\n    url=<span class=\"hljs-string\">\"https://news.ycombinator.com\"</span>,\n    query=<span class=\"hljs-string\">\"Extract each story: title, url, score, author\"</span>,\n    usage=usage,\n)\n\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Prompt tokens:     <span class=\"hljs-subst\">{usage.prompt_tokens}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Completion tokens: <span class=\"hljs-subst\">{usage.completion_tokens}</span>\"</span>)\n<span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"Total tokens:      <span class=\"hljs-subst\">{usage.total_tokens}</span>\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>The <code>usage</code> parameter is optional — omitting it changes nothing (fully backward-compatible). You can also reuse the same accumulator across multiple calls to get a grand total:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\">usage = TokenUsage()\nschema1 = JsonCssExtractionStrategy.generate_schema(url=url1, query=q1, usage=usage)\nschema2 = JsonCssExtractionStrategy.generate_schema(url=url2, query=q2, usage=usage)\nprint(f<span class=\"hljs-string\">\"Grand total: {usage.total_tokens} tokens\"</span>)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<p>Both <code>generate_schema</code> (sync) and <code>agenerate_schema</code> (async) support the <code>usage</code> parameter.</p>\n<h3 id=\"llm-provider-options\">LLM Provider Options</h3>\n<ol>\n<li><strong>OpenAI GPT-4 (<code>openai/gpt4o</code>)</strong></li>\n<li>Default provider</li>\n<li>Requires an API token</li>\n<li>Generally provides more accurate schemas</li>\n<li>\n<p>Set via environment variable: <code>OPENAI_API_KEY</code></p>\n</li>\n<li>\n<p><strong>Ollama (<code>ollama/llama3.3</code>)</strong></p>\n</li>\n<li>Open source alternative</li>\n<li>No API token required</li>\n<li>Self-hosted option</li>\n<li>Good for development and testing</li>\n</ol>\n<h3 id=\"benefits-of-schema-generation\">Benefits of Schema Generation</h3>\n<ol>\n<li><strong>One-Time Cost</strong>: While schema generation uses LLM, it's a one-time cost. The generated schema can be reused for unlimited extractions without further LLM calls.</li>\n<li><strong>Smart Pattern Recognition</strong>: The LLM analyzes the HTML structure and identifies common patterns, often producing more robust selectors than manual attempts.</li>\n<li><strong>Automatic Nesting</strong>: Complex nested structures are automatically detected and properly represented in the schema.</li>\n<li><strong>Learning Tool</strong>: The generated schemas serve as excellent examples for learning how to write your own schemas.</li>\n</ol>\n<h3 id=\"best-practices\">Best Practices</h3>\n<ol>\n<li><strong>Review Generated Schemas</strong>: While the generator is smart, always review and test the generated schema before using it in production.</li>\n<li><strong>Provide Representative HTML</strong>: The better your sample HTML represents the overall structure, the more accurate the generated schema will be.</li>\n<li><strong>Consider Both CSS and XPath</strong>: Try both schema types and choose the one that works best for your specific case.</li>\n<li><strong>Cache Generated Schemas</strong>: Since generation uses LLM, save successful schemas for reuse.</li>\n<li><strong>API Token Security</strong>: Never hardcode API tokens. Use environment variables or secure configuration management.</li>\n<li><strong>Choose Provider Wisely</strong>:</li>\n<li>Use OpenAI for production-quality schemas</li>\n<li>Use Ollama for development, testing, or when you need a self-hosted solution</li>\n</ol>\n<h3 id=\"multi-sample-schema-generation\">Multi-Sample Schema Generation</h3>\n<p>When scraping multiple pages with varying DOM structures (e.g., product pages where table rows appear in different positions), single-sample schema generation may produce <strong>fragile selectors</strong> like <code>tr:nth-child(6)</code> that break on other pages.</p>\n<p><strong>The Problem:</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-less\"><span class=\"hljs-selector-tag\">Page</span> <span class=\"hljs-selector-tag\">A</span>: <span class=\"hljs-selector-tag\">Manufacturer</span> <span class=\"hljs-selector-tag\">is</span> <span class=\"hljs-selector-tag\">in</span> <span class=\"hljs-selector-tag\">row</span> <span class=\"hljs-number\">6</span>  → <span class=\"hljs-selector-tag\">selector</span>: <span class=\"hljs-selector-tag\">tr</span><span class=\"hljs-selector-pseudo\">:nth-child</span>(<span class=\"hljs-number\">6</span>) <span class=\"hljs-selector-tag\">td</span> <span class=\"hljs-selector-tag\">a</span>\n<span class=\"hljs-selector-tag\">Page</span> <span class=\"hljs-selector-tag\">B</span>: <span class=\"hljs-selector-tag\">Manufacturer</span> <span class=\"hljs-selector-tag\">is</span> <span class=\"hljs-selector-tag\">in</span> <span class=\"hljs-selector-tag\">row</span> <span class=\"hljs-number\">5</span>  → <span class=\"hljs-selector-tag\">selector</span> <span class=\"hljs-selector-tag\">FAILS</span>\n<span class=\"hljs-selector-tag\">Page</span> <span class=\"hljs-selector-tag\">C</span>: <span class=\"hljs-selector-tag\">Manufacturer</span> <span class=\"hljs-selector-tag\">is</span> <span class=\"hljs-selector-tag\">in</span> <span class=\"hljs-selector-tag\">row</span> <span class=\"hljs-number\">7</span>  → <span class=\"hljs-selector-tag\">selector</span> <span class=\"hljs-selector-tag\">FAILS</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>The Solution:</strong> Provide multiple HTML samples so the LLM identifies stable patterns that work across all pages.</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-php-template\"><span class=\"language-xml\">from crawl4ai import JsonCssExtractionStrategy, LLMConfig\n\n# Collect HTML samples from different pages\nhtml_sample_1 = \"\"\"\n<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">table</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"specs\"</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Brand<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Apple<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Manufacturer<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">a</span> <span class=\"hljs-attr\">href</span>=<span class=\"hljs-string\">\"/m/apple\"</span>&gt;</span>Apple Inc<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">a</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">table</span>&gt;</span>\n\"\"\"\n\nhtml_sample_2 = \"\"\"\n<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">table</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"specs\"</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Manufacturer<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">a</span> <span class=\"hljs-attr\">href</span>=<span class=\"hljs-string\">\"/m/samsung\"</span>&gt;</span>Samsung<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">a</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Brand<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Galaxy<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">table</span>&gt;</span>\n\"\"\"\n\nhtml_sample_3 = \"\"\"\n<span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">table</span> <span class=\"hljs-attr\">class</span>=<span class=\"hljs-string\">\"specs\"</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Model<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Pixel 8<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Brand<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Google<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n  <span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">tr</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span>Manufacturer<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;<span class=\"hljs-name\">a</span> <span class=\"hljs-attr\">href</span>=<span class=\"hljs-string\">\"/m/google\"</span>&gt;</span>Google LLC<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">a</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">td</span>&gt;</span><span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">tr</span>&gt;</span>\n<span class=\"hljs-tag\">&lt;/<span class=\"hljs-name\">table</span>&gt;</span>\n\"\"\"\n\n# Combine samples with labels\ncombined_html = \"\"\"\n## HTML Sample 1 (Product A):\n```html\n\"\"\" + html_sample_1 + \"\"\"\n</span></code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"html-sample-2-product-b\">HTML Sample 2 (Product B):</h2>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-string\">\"\"\" + html_sample_2 + \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"html-sample-3-product-c\">HTML Sample 3 (Product C):</h2>\n<p></p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-string\">\"\"\" + html_sample_3 + \"\"\"</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n\"\"\"<p></p>\n<h1 id=\"provide-instructions-for-stable-selectors\">Provide instructions for stable selectors</h1>\n<p>query = \"\"\"\nIMPORTANT: I'm providing 3 HTML samples from different product pages.\nThe manufacturer field appears in different row positions across pages.\nGenerate selectors using stable attributes like href patterns (e.g., a[href*='/m/'])\ninstead of fragile positional selectors like nth-child().\nExtract: manufacturer name and link.\n\"\"\"</p>\n<h1 id=\"generate-schema-with-multi-sample-awareness\">Generate schema with multi-sample awareness</h1>\n<p>schema = JsonCssExtractionStrategy.generate_schema(\n    html=combined_html,\n    query=query,\n    schema_type=\"CSS\",\n    llm_config=LLMConfig(provider=\"openai/gpt-4o\", api_token=\"your-token\")\n)</p>\n<h1 id=\"the-generated-schema-will-use-stable-selectors-like\">The generated schema will use stable selectors like:</h1>\n<h1 id=\"ahrefm-instead-of-trnth-child6-td-a\">a[href*=\"/m/\"] instead of tr:nth-child(6) td a</h1>\n<p>print(schema)\n```</p>\n<p><strong>Key Points for Multi-Sample Queries:</strong></p>\n<ol>\n<li><strong>Format samples clearly</strong> - Use markdown headers and code blocks to separate samples</li>\n<li><strong>State the number of samples</strong> - \"I'm providing 3 HTML samples...\"</li>\n<li><strong>Explain the variation</strong> - \"...the manufacturer field appears in different row positions\"</li>\n<li><strong>Request stable selectors</strong> - \"Use href patterns, data attributes, or class names instead of nth-child\"</li>\n</ol>\n<p><strong>Stable vs Fragile Selectors:</strong></p>\n<table class=\"table table-striped table-hover\">\n<thead>\n<tr>\n<th>Fragile (single sample)</th>\n<th>Stable (multi-sample)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>tr:nth-child(6) td a</code></td>\n<td><code>a[href*=\"/m/\"]</code></td>\n</tr>\n<tr>\n<td><code>div:nth-child(3) .price</code></td>\n<td><code>.price, [data-price]</code></td>\n</tr>\n<tr>\n<td><code>ul li:first-child</code></td>\n<td><code>li[data-featured=\"true\"]</code></td>\n</tr>\n</tbody>\n</table>\n<p>This approach lets you generate schemas once that work reliably across hundreds of similar pages with varying structures.</p>\n<hr>\n<h2 id=\"11-conclusion\">11. Conclusion</h2>\n<p>With Crawl4AI's LLM-free extraction strategies - <code>JsonCssExtractionStrategy</code>, <code>JsonXPathExtractionStrategy</code>, and now <code>RegexExtractionStrategy</code> - you can build powerful pipelines that:</p>\n<ul>\n<li>Scrape any consistent site for structured data.  </li>\n<li>Support nested objects, repeating lists, or pattern-based extraction.  </li>\n<li>Scale to thousands of pages quickly and reliably.</li>\n</ul>\n<p><strong>Choosing the Right Strategy</strong>:</p>\n<ul>\n<li>Use <strong><code>RegexExtractionStrategy</code></strong> for fast extraction of common data types like emails, phones, URLs, dates, etc.</li>\n<li>Use <strong><code>JsonCssExtractionStrategy</code></strong> or <strong><code>JsonXPathExtractionStrategy</code></strong> for structured data with clear HTML patterns</li>\n<li>If you need both: first extract structured data with JSON strategies, then use regex on specific fields</li>\n</ul>\n<p><strong>Remember</strong>: For repeated, structured data, you don't need to pay for or wait on an LLM. Well-crafted schemas and regex patterns get you the data faster, cleaner, and cheaper—<strong>the real power</strong> of Crawl4AI.</p>\n<p><strong>Last Updated</strong>: 2025-05-02</p>\n<hr>\n<p>That's it for <strong>Extracting JSON (No LLM)</strong>! You've seen how schema-based approaches (either CSS or XPath) and regex patterns can handle everything from simple lists to deeply nested product catalogs—instantly, with minimal overhead. Enjoy building robust scrapers that produce consistent, structured JSON for your data pipelines!</p>\n</section>\n\n            <div class=\"page-actions-wrapper\"><button class=\"page-actions-button\" aria-label=\"Page copy\" aria-expanded=\"false\"><span>Page Copy</span></button><div class=\"page-actions-dropdown\" role=\"menu\">\n            <div class=\"page-actions-header\">Page Copy</div>\n            <ul class=\"page-actions-menu\">\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link\" id=\"action-copy-markdown\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-copy\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Copy as Markdown</span>\n                            <span class=\"page-action-description\">Copy page for LLMs</span>\n                        </span>\n                    </a>\n                </li>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-view-markdown\" target=\"_blank\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-view\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">View as Markdown</span>\n                            <span class=\"page-action-description\">Open raw source</span>\n                        </span>\n                    </a>\n                </li>\n                <div class=\"page-actions-divider\"></div>\n                <li class=\"page-action-item\">\n                    <a href=\"#\" class=\"page-action-link page-action-external\" id=\"action-open-chatgpt\" role=\"menuitem\">\n                        <span class=\"page-action-icon icon-ai\"></span>\n                        <span class=\"page-action-text\">\n                            <span class=\"page-action-label\">Open in ChatGPT</span>\n                            <span class=\"page-action-description\">Ask questions about this page</span>\n                        </span>\n                    </a>\n                </li>\n            </ul>\n            <div class=\"page-actions-footer\">ESC to close</div>\n        </div></div>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/migration/table_extraction_v073/",
    "title": "Migration Guide: Table Extraction v0.7.3 - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"migration-guide-table-extraction-v073\">Migration Guide: Table Extraction v0.7.3</h1>\n<h2 id=\"overview\">Overview</h2>\n<p>Version 0.7.3 introduces the <strong>Table Extraction Strategy Pattern</strong>, providing a more flexible and extensible approach to table extraction while maintaining full backward compatibility.</p>\n<h2 id=\"whats-new\">What's New</h2>\n<h3 id=\"strategy-pattern-implementation\">Strategy Pattern Implementation</h3>\n<p>Table extraction now follows the same strategy pattern used throughout Crawl4AI:</p>\n<ul>\n<li><strong>Consistent Architecture</strong>: Aligns with extraction, chunking, and markdown strategies</li>\n<li><strong>Extensibility</strong>: Easy to create custom table extraction strategies</li>\n<li><strong>Better Separation</strong>: Table logic moved from content scraping to dedicated module</li>\n<li><strong>Full Control</strong>: Fine-grained control over table detection and extraction</li>\n</ul>\n<h3 id=\"new-classes\">New Classes</h3>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    TableExtractionStrategy,    <span class=\"hljs-comment\"># Abstract base class</span>\n    DefaultTableExtraction,      <span class=\"hljs-comment\"># Current implementation (default)</span>\n    NoTableExtraction           <span class=\"hljs-comment\"># Explicitly disable extraction</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"backward-compatibility\">Backward Compatibility</h2>\n<p><strong>✅ All existing code continues to work without changes.</strong></p>\n<h3 id=\"no-changes-required\">No Changes Required</h3>\n<p>If your code looks like this, it will continue to work:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># This still works exactly the same</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=7\n)\nresult = await crawler.arun(url, config)\ntables = result.tables  <span class=\"hljs-comment\"># Same structure, same data</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"what-happens-behind-the-scenes\">What Happens Behind the Scenes</h3>\n<p>When you don't specify a <code>table_extraction</code> strategy:</p>\n<ol>\n<li><code>CrawlerRunConfig</code> automatically creates <code>DefaultTableExtraction</code></li>\n<li>It uses your <code>table_score_threshold</code> parameter</li>\n<li>Tables are extracted exactly as before</li>\n<li>Results appear in <code>result.tables</code> with the same structure</li>\n</ol>\n<h2 id=\"new-capabilities\">New Capabilities</h2>\n<h3 id=\"1-explicit-strategy-configuration\">1. Explicit Strategy Configuration</h3>\n<p>You can now explicitly configure table extraction:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-sql\"># <span class=\"hljs-keyword\">New</span>: Explicit control\nstrategy <span class=\"hljs-operator\">=</span> DefaultTableExtraction(\n    table_score_threshold<span class=\"hljs-operator\">=</span><span class=\"hljs-number\">7</span>,\n    min_rows<span class=\"hljs-operator\">=</span><span class=\"hljs-number\">2</span>,              # <span class=\"hljs-keyword\">New</span>: minimum <span class=\"hljs-type\">row</span> <span class=\"hljs-keyword\">filter</span>\n    min_cols<span class=\"hljs-operator\">=</span><span class=\"hljs-number\">2</span>,              # <span class=\"hljs-keyword\">New</span>: minimum <span class=\"hljs-keyword\">column</span> <span class=\"hljs-keyword\">filter</span>\n    verbose<span class=\"hljs-operator\">=</span><span class=\"hljs-literal\">True</span>             # <span class=\"hljs-keyword\">New</span>: detailed logging\n)\n\nconfig <span class=\"hljs-operator\">=</span> CrawlerRunConfig(\n    table_extraction<span class=\"hljs-operator\">=</span>strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"2-disable-table-extraction\">2. Disable Table Extraction</h3>\n<p>Improve performance when tables aren't needed:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># New: Skip table extraction entirely</span>\nconfig = CrawlerRunConfig(\n    table_extraction=NoTableExtraction()\n)\n<span class=\"hljs-comment\"># No CPU cycles spent on table detection/extraction</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"3-custom-extraction-strategies\">3. Custom Extraction Strategies</h3>\n<p>Create specialized extractors:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ruby\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">MyTableExtractor</span>(<span class=\"hljs-title class_\">TableExtractionStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_tables</span>(<span class=\"hljs-params\"><span class=\"hljs-variable language_\">self</span>, element, **kwargs</span>):\n        <span class=\"hljs-comment\"># Custom extraction logic</span>\n        <span class=\"hljs-keyword\">return</span> custom_tables\n\nconfig = <span class=\"hljs-title class_\">CrawlerRunConfig</span>(\n    table_extraction=<span class=\"hljs-title class_\">MyTableExtractor</span>()\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"migration-scenarios\">Migration Scenarios</h2>\n<h3 id=\"scenario-1-basic-usage-no-changes-needed\">Scenario 1: Basic Usage (No Changes Needed)</h3>\n<p><strong>Before (v0.7.2):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-built_in\">config</span> = CrawlerRunConfig()\nresult = await crawler.arun(url, <span class=\"hljs-built_in\">config</span>)\n<span class=\"hljs-keyword\">for</span> <span class=\"hljs-built_in\">table</span> <span class=\"hljs-keyword\">in</span> result.tables:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">table</span>[<span class=\"hljs-string\">'headers'</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>After (v0.7.3):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"># Exactly the same - no changes required\n<span class=\"hljs-built_in\">config</span> = CrawlerRunConfig()\nresult = await crawler.arun(url, <span class=\"hljs-built_in\">config</span>)\n<span class=\"hljs-keyword\">for</span> <span class=\"hljs-built_in\">table</span> <span class=\"hljs-keyword\">in</span> result.tables:\n    <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-built_in\">table</span>[<span class=\"hljs-string\">'headers'</span>])\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"scenario-2-custom-threshold-no-changes-needed\">Scenario 2: Custom Threshold (No Changes Needed)</h3>\n<p><strong>Before (v0.7.2):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-lua\"><span class=\"hljs-built_in\">config</span> = CrawlerRunConfig(\n    table_score_threshold=<span class=\"hljs-number\">5</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>After (v0.7.3):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Still works the same</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=5\n)\n\n<span class=\"hljs-comment\"># Or use new explicit approach for more control</span>\nstrategy = DefaultTableExtraction(\n    table_score_threshold=5,\n    min_rows=2  <span class=\"hljs-comment\"># Additional filtering</span>\n)\nconfig = CrawlerRunConfig(\n    table_extraction=strategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"scenario-3-advanced-filtering-new-feature\">Scenario 3: Advanced Filtering (New Feature)</h3>\n<p><strong>Before (v0.7.2):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Had to filter after extraction</span>\nconfig = CrawlerRunConfig(\n    table_score_threshold=<span class=\"hljs-number\">5</span>\n)\nresult = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config)\n\n<span class=\"hljs-comment\"># Manual filtering</span>\nlarge_tables = [\n    t <span class=\"hljs-keyword\">for</span> t <span class=\"hljs-keyword\">in</span> result.tables \n    <span class=\"hljs-keyword\">if</span> <span class=\"hljs-built_in\">len</span>(t[<span class=\"hljs-string\">'rows'</span>]) &gt;= <span class=\"hljs-number\">5</span> <span class=\"hljs-keyword\">and</span> <span class=\"hljs-built_in\">len</span>(t[<span class=\"hljs-string\">'headers'</span>]) &gt;= <span class=\"hljs-number\">3</span>\n]\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>After (v0.7.3):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Filter during extraction (more efficient)</span>\nstrategy = DefaultTableExtraction(\n    table_score_threshold=5,\n    min_rows=5,\n    min_cols=3\n)\nconfig = CrawlerRunConfig(\n    table_extraction=strategy\n)\nresult = await crawler.arun(url, config)\n<span class=\"hljs-comment\"># result.tables already filtered</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"code-organization-changes\">Code Organization Changes</h2>\n<h3 id=\"module-structure\">Module Structure</h3>\n<p><strong>Before (v0.7.2):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-markdown\">crawl4ai/\n  content<span class=\"hljs-emphasis\">_scraping_</span>strategy.py\n<span class=\"hljs-bullet\">    -</span> LXMLWebScrapingStrategy\n<span class=\"hljs-bullet\">      -</span> is<span class=\"hljs-emphasis\">_data_</span>table()      # Table detection\n<span class=\"hljs-bullet\">      -</span> extract<span class=\"hljs-emphasis\">_table_</span>data() # Table extraction\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<p><strong>After (v0.7.3):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-bash\">crawl4ai/\n  content_scraping_strategy.py\n    - LXMLWebScrapingStrategy\n      <span class=\"hljs-comment\"># Table methods removed, uses strategy</span>\n\n  table_extraction.py (NEW)\n    - TableExtractionStrategy    <span class=\"hljs-comment\"># Base class</span>\n    - DefaultTableExtraction      <span class=\"hljs-comment\"># Moved logic here</span>\n    - NoTableExtraction          <span class=\"hljs-comment\"># New option</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"import-changes\">Import Changes</h3>\n<p><strong>New imports available (optional):</strong>\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># These are now available but not required for existing code</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    TableExtractionStrategy,\n    DefaultTableExtraction,\n    NoTableExtraction\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"performance-implications\">Performance Implications</h2>\n<h3 id=\"no-performance-impact\">No Performance Impact</h3>\n<p>For existing code, performance remains identical:\n- Same extraction logic\n- Same scoring algorithm\n- Same processing time</p>\n<h3 id=\"performance-improvements-available\">Performance Improvements Available</h3>\n<p>New options for better performance:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Skip tables entirely (faster)</span>\nconfig = CrawlerRunConfig(\n    table_extraction=NoTableExtraction()\n)\n\n<span class=\"hljs-comment\"># Process only specific areas (faster)</span>\nconfig = CrawlerRunConfig(\n    css_selector=<span class=\"hljs-string\">\"main.content\"</span>,\n    table_extraction=DefaultTableExtraction(\n        min_rows=5,  <span class=\"hljs-comment\"># Skip small tables</span>\n        min_cols=3\n    )\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"testing-your-migration\">Testing Your Migration</h2>\n<h3 id=\"verification-script\">Verification Script</h3>\n<p>Run this to verify your extraction still works:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">import</span> asyncio\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig\n\n<span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">verify_extraction</span>():\n    url = <span class=\"hljs-string\">\"your_url_here\"</span>\n\n    <span class=\"hljs-keyword\">async</span> <span class=\"hljs-keyword\">with</span> AsyncWebCrawler() <span class=\"hljs-keyword\">as</span> crawler:\n        <span class=\"hljs-comment\"># Test 1: Old approach</span>\n        config_old = CrawlerRunConfig(\n            table_score_threshold=<span class=\"hljs-number\">7</span>\n        )\n        result_old = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config_old)\n\n        <span class=\"hljs-comment\"># Test 2: New explicit approach</span>\n        <span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> DefaultTableExtraction\n        config_new = CrawlerRunConfig(\n            table_extraction=DefaultTableExtraction(\n                table_score_threshold=<span class=\"hljs-number\">7</span>\n            )\n        )\n        result_new = <span class=\"hljs-keyword\">await</span> crawler.arun(url, config_new)\n\n        <span class=\"hljs-comment\"># Compare results</span>\n        <span class=\"hljs-keyword\">assert</span> <span class=\"hljs-built_in\">len</span>(result_old.tables) == <span class=\"hljs-built_in\">len</span>(result_new.tables)\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">f\"✓ Both approaches extracted <span class=\"hljs-subst\">{<span class=\"hljs-built_in\">len</span>(result_old.tables)}</span> tables\"</span>)\n\n        <span class=\"hljs-comment\"># Verify structure</span>\n        <span class=\"hljs-keyword\">for</span> old, new <span class=\"hljs-keyword\">in</span> <span class=\"hljs-built_in\">zip</span>(result_old.tables, result_new.tables):\n            <span class=\"hljs-keyword\">assert</span> old[<span class=\"hljs-string\">'headers'</span>] == new[<span class=\"hljs-string\">'headers'</span>]\n            <span class=\"hljs-keyword\">assert</span> old[<span class=\"hljs-string\">'rows'</span>] == new[<span class=\"hljs-string\">'rows'</span>]\n\n        <span class=\"hljs-built_in\">print</span>(<span class=\"hljs-string\">\"✓ Table content identical\"</span>)\n\nasyncio.run(verify_extraction())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"deprecation-notes\">Deprecation Notes</h2>\n<h3 id=\"no-deprecations\">No Deprecations</h3>\n<ul>\n<li>All existing parameters continue to work</li>\n<li><code>table_score_threshold</code> in <code>CrawlerRunConfig</code> is still supported</li>\n<li>No breaking changes</li>\n</ul>\n<h3 id=\"internal-changes-transparent-to-users\">Internal Changes (Transparent to Users)</h3>\n<ul>\n<li><code>LXMLWebScrapingStrategy.is_data_table()</code> - Moved to <code>DefaultTableExtraction</code></li>\n<li><code>LXMLWebScrapingStrategy.extract_table_data()</code> - Moved to <code>DefaultTableExtraction</code></li>\n</ul>\n<p>These methods were internal and not part of the public API.</p>\n<h2 id=\"benefits-of-upgrading\">Benefits of Upgrading</h2>\n<p>While not required, using the new pattern provides:</p>\n<ol>\n<li><strong>Better Control</strong>: Filter tables during extraction, not after</li>\n<li><strong>Performance Options</strong>: Skip extraction when not needed</li>\n<li><strong>Extensibility</strong>: Create custom extractors for specific needs</li>\n<li><strong>Consistency</strong>: Same pattern as other Crawl4AI strategies</li>\n<li><strong>Future-Proof</strong>: Ready for upcoming advanced strategies</li>\n</ol>\n<h2 id=\"troubleshooting\">Troubleshooting</h2>\n<h3 id=\"issue-different-number-of-tables\">Issue: Different Number of Tables</h3>\n<p><strong>Cause</strong>: Threshold or filtering differences</p>\n<p><strong>Solution</strong>: \n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-makefile\"><span class=\"hljs-comment\"># Ensure same threshold</span>\nstrategy = DefaultTableExtraction(\n    table_score_threshold=7,  <span class=\"hljs-comment\"># Match your old setting</span>\n    min_rows=0,               <span class=\"hljs-comment\"># No filtering (default)</span>\n    min_cols=0                <span class=\"hljs-comment\"># No filtering (default)</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"issue-import-errors\">Issue: Import Errors</h3>\n<p><strong>Cause</strong>: Using new classes without importing</p>\n<p><strong>Solution</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Add imports if using new features</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> (\n    DefaultTableExtraction,\n    NoTableExtraction,\n    TableExtractionStrategy\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h3 id=\"issue-custom-strategy-not-working\">Issue: Custom Strategy Not Working</h3>\n<p><strong>Cause</strong>: Incorrect method signature</p>\n<p><strong>Solution</strong>:\n</p><div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ruby\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">CustomExtractor</span>(<span class=\"hljs-title class_\">TableExtractionStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">extract_tables</span>(<span class=\"hljs-params\"><span class=\"hljs-variable language_\">self</span>, element, **kwargs</span>):  <span class=\"hljs-comment\"># Correct signature</span>\n        <span class=\"hljs-comment\"># Not: extract_tables(self, html)</span>\n        <span class=\"hljs-comment\"># Not: extract(self, element)</span>\n        <span class=\"hljs-keyword\">return</span> tables_list\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div><p></p>\n<h2 id=\"getting-help\">Getting Help</h2>\n<p>If you encounter issues:</p>\n<ol>\n<li>Check your <code>table_score_threshold</code> matches previous settings</li>\n<li>Verify imports if using new classes</li>\n<li>Enable verbose logging: <code>DefaultTableExtraction(verbose=True)</code></li>\n<li>Review the <a href=\"../../core/table_extraction/\">Table Extraction Documentation</a></li>\n<li>Check <a href=\"../examples/table_extraction_example.py\">examples</a></li>\n</ol>\n<h2 id=\"summary\">Summary</h2>\n<ul>\n<li>✅ <strong>Full backward compatibility</strong> - No code changes required</li>\n<li>✅ <strong>Same results</strong> - Identical extraction behavior by default</li>\n<li>✅ <strong>New options</strong> - Additional control when needed</li>\n<li>✅ <strong>Better architecture</strong> - Consistent with Crawl4AI patterns</li>\n<li>✅ <strong>Ready for future</strong> - Foundation for advanced strategies</li>\n</ul>\n<p>The migration to v0.7.3 is seamless with no required changes while providing new capabilities for those who need them.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  },
  {
    "url": "https://docs.crawl4ai.com/migration/webscraping-strategy-migration/",
    "title": "WebScrapingStrategy Migration Guide - Crawl4AI Documentation (v0.9.x)",
    "content": "<section id=\"mkdocs-terminal-content\">\n    <h1 id=\"webscrapingstrategy-migration-guide\">WebScrapingStrategy Migration Guide</h1>\n<h2 id=\"overview\">Overview</h2>\n<p>Crawl4AI has simplified its content scraping architecture. The BeautifulSoup-based <code>WebScrapingStrategy</code> has been deprecated in favor of the faster LXML-based implementation. However, <strong>no action is required</strong> - your existing code will continue to work.</p>\n<h2 id=\"what-changed\">What Changed?</h2>\n<ol>\n<li><strong><code>WebScrapingStrategy</code> is now an alias</strong> for <code>LXMLWebScrapingStrategy</code></li>\n<li><strong>The BeautifulSoup implementation has been removed</strong> (~1000 lines of redundant code)</li>\n<li><strong><code>LXMLWebScrapingStrategy</code> inherits directly</strong> from <code>ContentScrapingStrategy</code></li>\n<li><strong>Performance remains optimal</strong> with LXML as the sole implementation</li>\n</ol>\n<h2 id=\"backward-compatibility\">Backward Compatibility</h2>\n<p><strong>Your existing code continues to work without any changes:</strong></p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># This still works perfectly</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> AsyncWebCrawler, CrawlerRunConfig, WebScrapingStrategy\n\nconfig = CrawlerRunConfig(\n    scraping_strategy=WebScrapingStrategy()  <span class=\"hljs-comment\"># Works as before</span>\n)\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"migration-options\">Migration Options</h2>\n<p>You have three options:</p>\n<h3 id=\"option-1-do-nothing-recommended\">Option 1: Do Nothing (Recommended)</h3>\n<p>Your code will continue to work. <code>WebScrapingStrategy</code> is permanently aliased to <code>LXMLWebScrapingStrategy</code>.</p>\n<h3 id=\"option-2-update-imports-optional\">Option 2: Update Imports (Optional)</h3>\n<p>For clarity, you can update your imports:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-comment\"># Old (still works)</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> WebScrapingStrategy\nstrategy = WebScrapingStrategy()\n\n<span class=\"hljs-comment\"># New (more explicit)</span>\n<span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> LXMLWebScrapingStrategy\nstrategy = LXMLWebScrapingStrategy()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h3 id=\"option-3-use-default-configuration\">Option 3: Use Default Configuration</h3>\n<p>Since <code>LXMLWebScrapingStrategy</code> is the default, you can omit the strategy parameter:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ini\"><span class=\"hljs-comment\"># Simplest approach - uses LXMLWebScrapingStrategy by default</span>\n<span class=\"hljs-attr\">config</span> = CrawlerRunConfig()\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"type-hints\">Type Hints</h2>\n<p>If you use type hints, both work:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-python\"><span class=\"hljs-keyword\">from</span> crawl4ai <span class=\"hljs-keyword\">import</span> WebScrapingStrategy, LXMLWebScrapingStrategy\n\n<span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">process_with_strategy</span>(<span class=\"hljs-params\">strategy: WebScrapingStrategy</span>) -&gt; <span class=\"hljs-literal\">None</span>:\n    <span class=\"hljs-comment\"># Works with both WebScrapingStrategy and LXMLWebScrapingStrategy</span>\n    <span class=\"hljs-keyword\">pass</span>\n\n<span class=\"hljs-comment\"># Both are valid</span>\nprocess_with_strategy(WebScrapingStrategy())\nprocess_with_strategy(LXMLWebScrapingStrategy())\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"subclassing\">Subclassing</h2>\n<p>If you've subclassed <code>WebScrapingStrategy</code>, it continues to work:</p>\n<div class=\"highlight\"><pre><span></span><code data-highlighted=\"yes\" class=\"hljs language-ruby\"><span class=\"hljs-keyword\">class</span> <span class=\"hljs-title class_\">MyCustomStrategy</span>(<span class=\"hljs-title class_\">WebScrapingStrategy</span>):\n    <span class=\"hljs-keyword\">def</span> <span class=\"hljs-title function_\">__init__</span>(<span class=\"hljs-params\"><span class=\"hljs-variable language_\">self</span></span>):\n        <span class=\"hljs-variable language_\">super</span>().__init__()\n        <span class=\"hljs-comment\"># Your custom code</span>\n</code><button class=\"copy-code-button\" type=\"button\" aria-label=\"Copy code to clipboard\" title=\"Copy code to clipboard\">Copy</button></pre></div>\n<h2 id=\"performance-benefits\">Performance Benefits</h2>\n<p>By consolidating to LXML:\n- <strong>10-20x faster</strong> HTML parsing for large documents\n- <strong>Lower memory usage</strong>\n- <strong>Consistent behavior</strong> across all use cases\n- <strong>Simplified maintenance</strong> and bug fixes</p>\n<h2 id=\"summary\">Summary</h2>\n<p>This change simplifies Crawl4AI's internals while maintaining 100% backward compatibility. Your existing code continues to work, and you get better performance automatically.</p>\n</section>",
    "success": true,
    "error": "",
    "crawled_at": 1786439010.7467635,
    "is_shell": false
  }
]