Search by

lbonnet / seo-bundle

lbonnet-gda

A Symfony bundle that crawls a site once to audit its links, on-page content and technical SEO signals: broken links, titles and descriptions, canonical tags, indexing directives, redirects, robots.txt, sitemaps and hreflang.

Package info

github.com/lbonnet-gda/seo-bundle

Type:symfony-bundle

pkg:composer/lbonnet/seo-bundle

Statistics

Installs: 3

Dependents: 0

Suggesters: 0

Stars: 0

Open Issues: 1

v0.1.1 2026-09-28 12:46 UTC

This package is auto-updated.

Last update: 2026-09-29 13:34:43 UTC


README

CI Latest Version PHP Version License

A Symfony bundle that crawls a site once and audits it on three fronts: its links, its on-page content, and the technical signals that decide whether a page gets indexed, and indexed once.

Designed to run outside the request/response cycle — as a console command, a scheduled cron, or an async Messenger worker — so it fits both CI pipelines and continuous monitoring of a live site.

Unlike a content crawler, it never lets the HTTP client follow redirects: a 3xx is a finding, not a detour. Every page is requested with max_redirects = 0 and chains are walked explicitly, so each hop stays visible.

Requirements

  • PHP >= 8.1
  • Symfony 6.4, 7.x, or 8.x

Installation

composer require lbonnet/seo-bundle

If you don't use Symfony Flex, enable the bundle manually in config/bundles.php:

return [
    // ...
    Lbonnet\SeoBundle\SeoBundle::class => ['all' => true],
];

Modules

The audit runs three modules over a single crawl. Each one can be turned off, and a module that is off sends no request:

Module What it audits
links Broken internal and external links
on_page Titles, meta descriptions, headings and images
technical Canonical tags, indexing directives, robots.txt, sitemaps, redirects, hreflang

Quick start

# config/packages/seo.yaml
seo:
    base_url: 'https://example.com'
php bin/console seo:check

The command prints every issue it finds and exits with 1 when one of them is an error, so it doubles as a CI check. The audit also runs as a Messenger message, on a schedule, or behind your own event listener.

Documentation

  • Checks — the 60 checks, their severity, and what each one catches
  • Configuration — every option, with its default
  • Usage — console command, Messenger, Scheduler, and notifications
  • Reports — the JSON written after each audit

Known trade-offs

  • Requests go out one at a time. The crawl and the external link checks are sequential, so a large audit takes as long as the network makes it. Concurrency is planned and needs non-blocking per-host throttling first.
  • A redirect target is requested twice: once while resolving the chain (headers only, the body is canceled), then again to read its markup. This keeps chain resolution independent of crawling, at the cost of one extra HEAD-sized request per redirect.
  • Pages are read with libxml, not with DomCrawler, whose parser changes across PHP and Symfony versions and, through masterminds/html5, never closes <head> early. Like browsers, libxml closes <head> on the usual culprits (a stray <div>, a tracking <img>, stray text, a misplaced <iframe>), but not on an <svg> or a custom element. It also keeps a <noscript> holding an <img> inside <head>, which is how a JavaScript-enabled crawler reads it. This is what canonical_not_in_head and hreflang_not_in_head rely on.
  • JavaScript is not executed. A link, a canonical, or a title that only exists after hydration is invisible to the audit, as it is to a search engine that does not render the page.

Security

To report a vulnerability, please don't open a public issue — see SECURITY.md for how to report it privately.

License

MIT — see LICENSE.