Search by

bitandblack / document-crawler

TobiasKöngeterBit&Black

Extract titles, meta tags, images, anchors, headings, canonical URLs, structured data and more from HTML or XML documents, based on Symfony's DomCrawler.

Package info

github.com/BitAndBlack/document-crawler

pkg:composer/bitandblack/document-crawler

Statistics

Installs: 1 001

Dependents: 0

Suggesters: 0

Stars: 7

Open Issues: 0

0.7.0 2026-10-09 06:37 UTC

This package is auto-updated.

Last update: 2026-10-09 06:38:20 UTC


README

PHP from Packagist Latest Stable Version Total Downloads License

Bit&Black Logo

Bit&Black Document Crawler

Extract titles, meta tags, images, anchors, headings, canonical URLs, structured data and more from an HTML or XML document.

Installation

This library is installed via Composer:

composer require bitandblack/document-crawler

It requires PHP 8.2 or higher.

Downloading resources or using HolisticDocumentCrawler::createFromUrl() additionally requires a PSR-18 HTTP client implementation in your project (for example symfony/http-client) — the HttpDiscoveryClient picks up whatever is installed.

Usage

Using crawlers to extract parts of a document

The Bit&Black Document Crawler library provides different crawlers to extract information from a document. The following crawlers are currently available:

  • AnchorsCrawler: Crawl and extract all defined anchors in a document, that have been declared with <a href="...">...</a>.
  • CanonicalCrawler: Crawl and extract the canonical URL of a document, that has been declared with <link rel="canonical" href="..." />.
  • HeadingsCrawler: Crawl and extract all headings in a document, that have been declared with <h1>...</h1> up to <h6>...</h6>.
  • IconsCrawler: Crawl and extract all defined icons in a document, that have been declared with <link rel="icon" ... />.
  • IframesCrawler: Crawl and extract all defined iframes in a document, that have been declared with <iframe ...></iframe>.
  • ImagesCrawler: Crawl and extract all defined images in a document, that have been declared with <img ... />.
  • LanguageCodeCrawler: Crawl and extract the language code of a document, that has been declared with <html lang="...">.
  • LinkTagsCrawler: Crawl and extract all link tags of a document, that have been declared with <link ... />.
  • MetaTagsCrawler: Crawl and extract all defined meta tags in a document, that have been declared with <meta ... />.
  • StructuredDataCrawler: Crawl and extract all structured data blocks in a document, that have been declared with <script type="application/ld+json">...</script>.
  • TitleCrawler: Crawl and extract the title of a document, that has been declared with <title>...</title>.

All those crawlers work the same — they need a DomCrawler object, that contains the document:

<?php

use BitAndBlack\DocumentCrawler\Crawler\TitleCrawler;
use Symfony\Component\DomCrawler\Crawler;

$document = <<<HTML
<!doctype html>
<html lang="en">
    <head>
        <title>Test</title>
    </head>
    <body>
        <h1>Hello world</h1>
    </body>
</html>
HTML;

$crawler = new Crawler($document);

$titleCrawler = new TitleCrawler($crawler);
$titleCrawler->crawlContent();

// This will output `Test`.
echo $titleCrawler->getTitle();

You can create a custom Crawler by implementing the CrawlerInterface.

All DTOs implement JsonSerializable and Stringable, so extracted results can be encoded or cast to string directly.

Handling resources

In some cases, crawlers process external resources, which you may want to handle in a specific way. To achieve this, each crawler uses a so-called Resource Handler. The following resource handlers are currently available:

  • The FileSystemDownloadHandler: This one loads resources and writes them to the file system. There are different Http Clients available to fetch resources:

    • The HttpDiscoveryClient is the default one and makes use of whatever library your project uses to download resources.
    • The ReactClient needs the react/http library and downloads resources asynchronously in the background: the downloads run in parallel and the returned download item reflects the final status of the download, once it has finished.
    • You can — for sure — create a custom Http Client by implementing the HttpClientInterface.
  • The PassiveResourceHandler: This handler does nothing and is the default one.

You can create a custom Resource Handler by implementing the ResourceHandlerInterface.

Crawling everything at once

In case you don't want to set up every crawler yourself, there is the HolisticDocumentCrawler, that does all the work for you:

<?php

use BitAndBlack\DocumentCrawler\HolisticDocumentCrawler;

$document = <<<HTML
<!doctype html>
<html lang="en">
    <head>
        <title>Test</title>
    </head>
    <body>
        <h1>Hello world</h1>
    </body>
</html>
HTML;

$holisticDocumentCrawler = new HolisticDocumentCrawler($document);

// Get all anchors:
$anchors = $holisticDocumentCrawler->getAnchors();

// Get the canonical URL:
$canonicalUrl = $holisticDocumentCrawler->getCanonicalUrl();

// Get all headings:
$headings = $holisticDocumentCrawler->getHeadings();

// Get all icons:
$icons = $holisticDocumentCrawler->getIcons();

// Get all iframes:
$iframes = $holisticDocumentCrawler->getIframes();

// Get all images:
$images = $holisticDocumentCrawler->getImages();

// Get the language code:
$languageCode = $holisticDocumentCrawler->getLanguageCode();

// Get all link tags:
$linkTags = $holisticDocumentCrawler->getLinkTags();

// Get all meta tags:
$metaTags = $holisticDocumentCrawler->getMetaTags();

// Get all structured data blocks:
$structuredData = $holisticDocumentCrawler->getStructuredData();

// Get the title:
$title = $holisticDocumentCrawler->getTitle();

The HolisticDocumentCrawler can also be initialised using the createFromUrl method:

<?php

use BitAndBlack\DocumentCrawler\HolisticDocumentCrawler;

$holisticDocumentCrawler = HolisticDocumentCrawler::createFromUrl('https://www.bitandblack.com');

Runnable examples can be found in the examples directory. The list of notable changes is documented in the CHANGELOG.md.

Help

If you have any questions, feel free to contact us at hello@bitandblack.com.

Further information about Bit&Black can be found under www.bitandblack.com.