Skip to content

Aggregate: headless, AI-ready web analytics ​

Web analytics that your BI tools and AI assistants can query, without identifying your visitors.

Aggregate counts what happens on your websites, not who did it. It keeps the results in your own database as stable SQL views with privacy protection built in, so Power BI, Tableau, Looker or an AI assistant can answer questions about your traffic directly. There is no analytics dashboard to learn, and the views a report or a model reads hold no visitor-level data.

  • Headless. Collection, configuration and maintenance run from YAML and the command line; an optional admin UI manages the installation. Reports live in the tools you already use.
  • AI-ready. Give an assistant the same read-only account as your BI tool. It sees documented, versioned views and a built-in data dictionary, never raw rows, identifiers or counts below your disclosure threshold.
  • Privacy enforced, not promised. The server stores no visitor ID, IP address or exact time in anonymous mode and adds detail only with consent, and the database withholds small counts from every report, extract and query.
  • Yours to run. Self-hosted from a release ZIP on ordinary PHP hosting, with MySQL, MariaDB, PostgreSQL, SQL Server or SQLite.

Beta software: contributors and early adopters are welcome. Start with a disposable installation and synthetic events. Help test installation, consent, headless administration, and database reporting through the beta testing guide. Review its known limitations before using the platform with live traffic; no independent security or privacy audit is claimed.

Aggregate logo

Aggregate's anonymous mode creates no visitor identifier. By default, an anonymous event holds a sanitized path, a coarse referrer channel, device and viewport buckets, and a UTC hour. It stores no visitor or session ID, IP address, User-Agent string, or exact event timestamp. Optional goals, coarse geography, organization markers, and explicitly allowlisted properties can add context; their values still need review for identifying detail.

Useful insight does not require a visit or session ID. Broad medium values, when explicitly allowed, can complement channels, page-path activity, event tracking and goal counts to help paint a picture of what is working. See the measurement guide for examples and current reporting boundaries, and the roadmap for planned work.

When you need more, enhanced mode adds identifiers, properties and exact dimensions — but only for visitors who have made an affirmative choice, and it stops the moment they withdraw it.

Once you start collecting data, connect Power BI, Tableau, Looker or an AI assistant to the approved reporting views.

Dashboard screenshot

What you get ​

Privacy-minimized event collection. Page views and named events, stored with sanitized paths, coarse dimensions and hour buckets. Emails, UUIDs and numeric route IDs are stripped out of paths before storage. Optional allowlisted conversion goals and continent- or country-level geography. Keep raw events private and use the thresholded reporting views for routine BI access.

Suppression in the view, not the dashboard. bi_anonymous_events_v1 and its siblings withhold the current bucket and hide any cell below your configured minimum. The threshold lives in the database and applies to every consumer of these views. Keep raw input tables private when granting a BI connection access.

Ready for AI analysis. The reporting views keep their names and column meanings across releases, and the glossary views label every code and describe every column: the context a model needs to write correct SQL instead of guessing. An assistant connected with the read-only reporting account gets the same protection as a BI user, and what reaches a model provider is released counts, not visitor records. See connect an AI assistant.

Consent that actually toggles something. setConsent(true) turns on visitor and session IDs, custom properties and exact dimensions. setConsent(false) drops the identifiers and returns to coarse rows. Wire it to your CMP and the behavior matches what the banner promised.

Optional administration UI. Operate collection, setup, and maintenance through YAML and CLI commands; set DASHBOARD_ENABLED=0 to disable dashboard/login routes. BI disclosure thresholds currently require the admin UI or controlled database administration; see configuration sources. Visualize analytics in your BI tool; the application UI manages the installation and its data contracts.

Rebrandable without a fork. Name, logo, colors and fonts are configuration. Contrast ratios are validated, so an unreadable palette gets rejected rather than shipped.

Portable infrastructure. PHP 8.2+ on PostgreSQL, MySQL, MariaDB, SQL Server or SQLite, with native and Docker deployment options. Size the host for your traffic, retention policy, and database; beta reports should include the workload tested.

A way to exclude your own team. Mark staff browsers with a configurable cookie or local storage flag and filter them out in your BI tool, without dropping the data or trusting an IP range.

Details:

  • Privacy-minimized collection: page views and safe named events, with optional allowlisted goals and coarse local geography.
  • Reporting views with suppression: completed hourly or daily aggregates with configurable minimum event counts.
  • AI-ready reporting: versioned views and a built-in data dictionary that AI assistants can query with the same read-only account as BI tools.
  • Optional page depth: a capped page count shared by events on each page, with tab storage or URL parameter passing configured through UI/YAML. See storage, URL and reporting boundaries before enabling it.
  • Consent-based enhanced detail: visitor/session IDs, properties, and exact dimensions when enabled by your consent manager.
  • Headless operation: an ingestion API, YAML configuration, and CLI commands, with an optional admin dashboard.
  • Server-side collection: send page views and goals from your backend, with no script on the page. See the server-side guide for PHP, Java, .NET, Node.js, Python, Ruby and Go examples.
  • Organization traffic markers: mark team browsers with a configurable cookie or local storage entry, defaulting to orgInternalTraffic=true, and filter retained event JSON in BI reports.
  • Configurable branding and lifecycle: dashboard name, logo, colors, fonts, archiving, and retention policies.
  • Database choice: PostgreSQL, MySQL, MariaDB, SQL Server, or SQLite. Enhanced ingestion uses Symfony Messenger with synchronous or asynchronous delivery.

What you give up ​

This matters more than the feature list, so it's up front.

No unique visitors, sessions or bounce rate in anonymous mode. Not because they haven't been built yet, but because computing them requires exactly the identifier this project refuses to create. Anonymous mode counts events. One person can contribute several of them to the same cell.

Small numbers disappear. A cell below your threshold is withheld, and widening the time window in your BI tool won't recover it. Low-traffic sites can have persistent gaps; combining periods only combines cells already released by the views.

Reports lag one bucket. The current UTC hour is never released for events; the current UTC day is never released for goals and geography. There is no real-time view.

No built-in assistant or charts. Aggregate makes the data safe and understandable to query. The BI tool, the model and the connector are yours to choose, and an assistant's access deserves the same review as any BI user's.

"Anonymous" is the name of a mode, not a legal conclusion. Rare paths, unusual event names, small populations and outside information can still make a row personal in context. Withdrawing consent is prospective — it stops future enhanced detail, it does not erase what the server already holds. The compliance guide is specific about where the line sits and what remains your responsibility.

This project suits teams that want self-hosted event measurement feeding their existing BI tools and AI assistants. It does not provide session replay, heatmaps, or a built-in marketing attribution suite.

Is this for you? ​

Likely yes if you answer to a privacy office or DPO, you run a public-sector, health, education or legal site, you already own a BI stack and want measurement to feed it rather than compete with it, you want to ask an AI assistant about your traffic without handing it visitor-level data, or you want to be able to explain your whole data model on one page.

Likely no if you need unique-visitor counts, funnels, session replay, heatmaps or marketing attribution. Plausible and Matomo are good software and will serve you better. This isn't trying to win that comparison.

Quick start ​

On a web host, from a release ZIP ​

No terminal is needed:

  1. Download aggregate-VERSION.zip from the Assets of the latest release (not the Source code archives).
  2. Upload and extract it with your hosting panel's file manager, and set the site's document root to the extracted public folder.
  3. Open the site. The setup page checks the server, asks for the one-time code in SETUP-CODE.txt, tests your database connection and writes .env.local. Then create the administrator account.

See Release ZIP on a web host for the details and step-by-step hosting panel guides (Plesk so far). PHP 8.2+ and MySQL 8.0+, MariaDB 10.6+, PostgreSQL 13+, SQL Server 2017+ or SQLite are required.

For development ​

For this development setup, install PHP 8.2+, Composer, the required PHP extensions and database driver, Docker Compose, and Make. The deployment guide covers native and production installation; CONTRIBUTING.md covers development prerequisites in detail.

From a fresh checkout:

bash
git clone --branch development https://github.com/Subschema-LLC/aggregate.git
cd aggregate
cp .env.dev .env
cp config/aggregate.yaml.example config/aggregate.yaml
composer install
make start-mysql
make migrate-mysql

The checked-in Compose setup serves http://localhost:9002. Set the active YAML environment's app_host to that URL for local integration snippets. Open /install to create a dashboard administrator, then register a website in the dashboard or through the CLI:

bash
docker compose exec php php bin/console app:create-website
curl http://localhost:9002/api/health

Use start-postgres / migrate-postgres or start-mariadb / migrate-mariadb for another Docker profile. The sample environment is for development; use your own secrets and connection settings in production. Existing installations should follow the upgrade guidance before applying historical privacy migrations.

Admins choose an update method on the Updates page: signed release ZIPs (recommended), the repository for Git clones (advanced), or deployed another way when a hosting panel's Git deployment (such as Plesk or cPanel) or a CI/CD pipeline deploys the code. Install updates with the Install update button or php bin/console app:updates:apply; migrations and cache rebuilds run with a maintenance page in place. For code deployed another way, the page shows the deployed commit and how far behind it is, and php bin/console app:updates:deployed runs the same post-deployment steps from your tool's deployment action. updates_method and updates_branch (default master) can also be set in YAML or with app:updates:method. Your .env.local, config/aggregate.yaml, website and tag settings, config/*.local.yaml overrides and var/ data are never overwritten. See the update guide and signed release packages.

The Feature flags admin page and YAML can disable Updates or hide its navigation entries. Updates stays enabled by default. See the feature flag guide for configuration and contributor examples.

Add tracking to a website ​

Websites and the Setup wizard generate window configuration, query-parameter, or tag-manager installation snippets. For a direct tracker installation, use the website token from your registration and your analytics host:

html
<script>
  window.Aggregate = {
    endpoint: 'https://analytics.example.com/api/receive',
    websiteToken: 'your-website-token',
    consent: false
  };
</script>
<script src="https://analytics.example.com/aggregate.js?min=1" defer referrerpolicy="no-referrer"></script>

The tracker sends a page view automatically. Once it has loaded, record a named event with window.Aggregate.emit('button_click'). Configure goals and review event properties before using them.

Tag-manager snippets install the container and enabled CMP. To load analytics through it, explicitly enable the manager and add the supplied tracker URL as a script action. Remove any separate tracker installation to avoid duplicate page views.

Connect your consent manager to window.Aggregate.setConsent(true) only after an affirmative choice, and call setConsent(false) for rejection or withdrawal. The tracking and GTM guide covers custom events, tag setup, consent wiring, and troubleshooting.

A dedicated Aggregate Google Tag Manager tag template is under development in a separate repository. Follow that repository for template progress and availability. The tracking guide includes GTM Custom HTML examples.

Configuration and reporting ​

SettingSource of truth
Database, transport, app secret, proxy trustSymfony environment files or server environment variables
Branding, collection controls, organization markers, lifecycle, dashboard toggleconfig/aggregate.yaml and supported environment overrides
Custom property model, UTM/query mappings, consent requirements, reporting columnsconfig/aggregate.yaml, managed through the Data model admin page or YAML
Website domains, allowed event sources, and public ingestion tokensconfig/websites.yaml, managed through the dashboard, YAML, or CLI creation options
Per-website CMP, tags, triggers, and variablesconfig/tag-manager/sites/<site-id>.yaml, managed through the Tag manager admin page or YAML
Conversion-goal definitionsconfig/goals.yaml, customized in config/goals.local.yaml
Navigation labels and linksconfig/navigation.yaml, customized in config/navigation.local.yaml
BI disclosure thresholdsanalytics_privacy_settings in the database; dashboard or controlled database administration

For headless deployments, set dashboard_enabled: false in the active YAML environment and DASHBOARD_ENABLED=0, then clear Symfony's cache. BI thresholds remain database settings; they are not mirrored in YAML. The configuration reference explains defaults, overrides, branding, goals, organization markers, and retention.

Routine BI connections should use approved views:

ViewReports
bi_anonymous_events_v1Hourly page views and named events
bi_anonymous_goals_v1Daily occurrences of configured goals
bi_anonymous_geo_events_v1Daily coarse geography with additional suppression

Keep raw events, archive tables, and unsuppressed operational views private. Organization markers are stored under their configured name in raw event JSON; the grouped views and archives omit that flag. Connect BI tools and AI assistants lists every column, safe query patterns, a connection checklist and how to set up an AI assistant. See the compliance guide for access and disclosure rules, organization traffic for filtering, and the database guide for connections and schema details.

The Data model page provides UTM/query mappings, configurable anonymous property whitelists, downloadable YAML, and UI/CLI regeneration of private custom reporting views. All UTMs require consent by default. Anonymous attribution should use at most a broad utm_medium; administrators can override this recommendation with the documented warning about more detailed values.

Generate and copy synthetic event examples from the saved model, or export them with php bin/console app:analytics:examples. Optional value types and separate numeric reporting columns support external calculations without changing existing text columns. The ecommerce recipe uses flat properties and integer minor units for money. Administration pages and configurable submenus keep each task focused.

Use the optional JavaScript build for minified tracker and drop-in scripts. The Setup wizard provides installation steps, button callouts, and per-website CMP/script copy or download. The optional simple tag manager stores each site's settings in YAML, serves its scripts remotely from the Aggregate installation, and maps each tag to a consent category or an explicit option to run without consent. Tags can load scripts or call an existing library method, triggered by page readiness, browser events, or data-layer events with named variable references.

Documentation ​

All guides are also published as a searchable site that needs no account: subschema-llc.github.io/aggregate. It is built from these files on master and describes the current release.

GuideContents
Why Aggregate existsThe problem, the principles, and what the project deliberately will not do
Architecture tourHow an event flows from a web page to a BI report, and where to make common changes
Design decisionsWhy the main choices were made, what they cost, and what would change them
GlossaryTerms used across the code, settings and documentation
ContributingBranching strategy, development setup, architecture, privacy invariants, tests, Make commands, and pull requests
Beta testingFirst test session, expected privacy/reporting behavior, known limitations, and feedback
Agent guideProduct ethos, architecture boundaries, privacy rules, and development expectations for coding agents
RoadmapPlanned work, available foundations, and current update limitations
Security policyPrivate vulnerability reporting and disclosure guidance
Code of conductCommunity expectations and reporting concerns
Public release preparationMaintainer reporting setup, Git history review, GitHub checks, and launch steps
Privacy and complianceMeasurement limits, consent, organization traffic, BI suppression, logging, retention, and operator checks
ConfigurationYAML settings, environment overrides, branding, goals, and lifecycle policy
Feature flagsYAML/admin controls, navigation visibility, and developer/contributor guidance
Data modelUTM/query mappings, anonymous property whitelists, model sharing, and custom reporting columns
Event examplesUI copy/download, headless JSON exports, typed properties, and a flat ecommerce recipe
JavaScript buildOptional minification, dynamic tracker configuration, and build verification
UpdatingChoosing release ZIP or repository updates, setting up a Git clone, switching methods, recovery and troubleshooting
Release packagesPublishing from master, signing keys, installable ZIPs, verification, and update groundwork
Tracking and GTMBrowser integration, custom events, tag-manager examples, and troubleshooting
Server-side collectionSending events from your backend without a script, what that means for consent, and examples in seven languages
GTM tag templateSeparate repository for the Google Tag Manager tag template, under development
DeploymentDocker/native setup, web servers, workers, production operations, and upgrades
Plesk deploymentPlesk steps for a release ZIP without SSH, Git-based setup, and worker options
DatabaseSupported engines, connection strings, migrations, and reporting schema
Connect BI tools and AI assistantsApproved reporting views and their columns, safe queries, a connection checklist for Power BI or Tableau, and AI assistant setup

Contributing ​

Create a working branch from the latest development, using a descriptive name such as feature/add-goal-validation, issue/123-fix-consent, docs/update-setup, or chore/update-dependencies. Changes move through pull requests in this order:

  1. Working branch → development: contributors submit changes for review and integration.
  2. development → uat: maintainers promote changes for user acceptance testing (UAT).
  3. uat → master: maintainers promote accepted changes for production release.

Keep contributor PRs targeted at development; uat and master receive the promotion PRs above. See the branching strategy for branch roles and PR targets, and the release guide for tagging and publishing from master.

Contributions should preserve the privacy invariants and favor simple, portable, secure, maintainable designs. Start with CONTRIBUTING.md.

You can also contribute without writing application code: try a fresh install, exercise your database engine, reproduce a reported bug with synthetic data, review keyboard access and mobile administration, or improve a confusing setup step. The roadmap lists the highest-priority work; the beta testing guide explains what makes a useful test report.

Contributions developed with AI coding agents are welcome. Bring your own expertise and judgment to the collaboration: provide project context, guide the agent's decisions, and review and test the result. Please submit changes you understand and can explain, including how they fit Aggregate's architecture and privacy goals. The contributor remains responsible for the work they submit.

Follow the code of conduct, review the roadmap before proposing substantial work, and use the security policy for private vulnerability reports.

License ​

The license split is:

  • Browser tracker: public/aggregate.js is licensed under BSD-3-Clause, with the full text in its header and js/LICENSE.txt. This file is also the tracker source; there is no separate build source. The configured script served at /aggregate.js carries the same BSD license.
  • Code of Conduct: CODE_OF_CONDUCT.md adapts Contributor Covenant 2.1 under CC BY 4.0, with source attribution and license links in that file.
  • Everything else in this project's first-party code and documentation: GNU AGPL version 3 only (AGPL-3.0-only), under LICENSE. This includes the server, dashboard, and other JavaScript; the tracker exception does not change their license.

Third-party dependencies and vendored assets retain their own licenses, including the vendored Bulma MIT notice.

Server and documentation licensed under AGPL-3.0; the tracker under BSD-3-Clause.