Checking access…

Built on Google Cloud
Omni Alpha

Hedge Fund
Data Intelligence

Behavioural intelligence signals derived from 180B+ monthly CDR events — mobility scores, engagement indices, and churn propensity — delivered to quantitative hedge funds as structured BigQuery feeds and REST APIs for systematic trading strategies.

82M+ Subscribers 180B CDRs / month 5 GCP pipeline stages BQ + REST API delivery
What's Inside an Alpha Signal Report
Omni Alpha — Hedge Fund Data Intelligence
Product Overview

Alpha Signal Report

Omni Alpha delivers per-ticker consumer intelligence reports — tracking how Verizon subscribers engage with a public company's digital product across 10 behavioural funnel stages. Signals are correlated against reported financial KPIs to give quantitative hedge funds a real-time window into consumer activity before earnings.

Sample Report
FLUT US
Flutter Entertainment inc
Funnel Stages
10
Alpha Points
6
Peak / Day
845K
Window
90d
to Jul 2026
Alpha Point = key predictive signal URLs
Daily Unique Visitors — All Funnel Stages
Apr 6 → Jul 4, 2026  ·  260 tracked URLs

Consumer Engagement Funnel

Avg active subscribers / day by funnel stage  ·  Q1 2026  ·  ★ = Alpha Point signal

New User Acquisition
1.1K
★
Login
132.8K
★
Open App
188.5K
★
Marketing
218.2K
—
User Engagement
400.2K
★
Commerce Intent
177.8K
★
Transactional Activity
49.0K
★
Loyalty Activity
5.0K
—
Churn Risk
11.8K
—
Unclassified
893.2K
—

KPI Classification — FLUT US

Accuracy = signal-to-actual QoQ correlation · Relevance = strategic importance to hedge fund clients

Financial KPIScopeAccuracyRelevance
Europe Revenue EU 4486
US Revenue US 9864
International Revenue INTL 3390
Total Revenue ALL 7471
Signal Breakdown
Funnel StageSignal URLsAlpha PtsAvg / Day
New User Acquisition711.1K
Login122132.8K
Open App245188.5K
Marketing260218.2K
User Engagement447400.2K
Commerce Intent212177.8K
Transactional Activity22449.0K
Churn Risk17011.8K
Total26021—

What's Inside an Alpha Signal Report

Each report covers 8 analytical sections — from raw engagement trends through to AI-generated narrative insights — updated on a rolling 90-day window per ticker.

Section 01
Daily Visitor Trend
90-day line chart of daily unique subscribers across all funnel stages. Toggle individual signal filters (New User, Login, Open App, Commerce, Transactional) to isolate trends.
Section 02
Signal Breakdown
Summary table of all tracked funnel stages — URL count per stage, Alpha Point designation, and average daily unique visitors. Distinguishes high-signal vs. background traffic.
Section 03
KPI Classification
Accuracy and relevance scores for each tracked financial KPI. Accuracy measures telco-to-actual QoQ correlation; relevance captures hedge fund strategic importance.
Section 04
KPI Actuals & Trend
Quarterly KPI actuals vs. Wall Street consensus vs. Eureka model estimates, charted alongside any telco signal metric overlay to visualise lead time and predictive accuracy.
Section 05
Signal Correlation
Heatmap of Signal Score (0–100) between all 10 funnel stages × 17 engagement metrics × each KPI. ≥85 = very strong predictive signal. Score penalises direction mismatches.
Section 06
QoQ Funnel Comparison
Quarter-over-quarter horizontal bar comparison across all funnel stages. Toggle between 17 metrics — active users, sessions, data volume, duration per user, hits per session.
Section 07
URL Ranking
Individual URLs ranked by any engagement metric — identifies highest-traffic acquisition pages, commerce touchpoints, and transactional confirmation flows within the product.
Section 08
AI Insights
LLM-generated narrative synthesising signal trends, QoQ changes, and KPI outlook — delivered as actionable hedge fund commentary derived from all 7 preceding sections.

Delivery

Reports are delivered as interactive HTML — self-contained, no login required per file, shareable with portfolio managers and analysts.

📄
HTML Report
Self-contained interactive report with live charts, filters, and correlation heatmaps. Signed URL delivery via GCS.
🔌
BigQuery Feed
Raw signal tables shared in BQ dataset. Clients run their own queries against ticker × stage × date partitions.
🌐
REST API
JSON endpoints for programmatic access to signal scores, KPI actuals, and correlation data. Rate-limited at 10K req/min.
Data Requirement

Input Data Specifications

Four source dataset families delivered via encrypted GCS transfer. Two are generated daily at high velocity (web traffic sessions and location pings), two are refreshed monthly (subscriber demographics and reference master tables). Field widths and byte measurements are empirically validated from source file analysis.

Raw Input / day
~95 GB
Parquet compressed
Timestamped
35 TB/yr
Archive cold tier
5-min Agg
1.6 TB/yr
Nearline warm tier
15-min Agg
620 GB/yr
Standard hot tier
~60M subs
3 sources
Web · Location · Demo
Source 01
Web Traffic Events
Primary · Daily

Every subscriber HTTP/HTTPS session routed through the operator network. Each row is one TCP session: a unique pairing of a privacy-preserving subscriber token, a radio cell site, a destination hostname, and the data volumes transferred. The highest-volume source — approximately 1.25 billion raw session rows per day before any aggregation. Delivered as Parquet via encrypted GCS transfer; Pub/Sub triggers downstream Dataflow processing within 15 minutes of file landing.

Raw rows/day
~1.25 B
Raw bytes/row
~151 B
Compressed/row
~49 B
Parquet + Snappy 3.08×
Daily (compressed)
~58 GB
~21 TB/year
FieldTypeDescription
TOKEN_ID
VARCHAR(64)
Privacy-preserving hashed subscriber identifier. SHA-256 of MSISDN + daily salt. Consistent across sessions for the same device within a calendar day; re-salted at midnight UTC. Maps to the subscriber demographics table via an internal linkage layer — raw phone number is never stored at this level.
START_VISIT_DATE_TIME
TIMESTAMP TZ
Session start timestamp in UTC with timezone offset. Precision to the minute. Used as partition basis — event_date is extracted from this field. Buckets into 5-minute and 15-minute windows in the aggregated downstream tables.
CELL_ID
VARCHAR(20)
Unique identifier for the serving radio cell. Encodes site code and commissioning year: format {site_code}-{year}. Foreign key to the cell site reference table. Derives network-level geographic location without storing raw lat/lon at this table level.
VALID_FROM_CELL
TIMESTAMP
Timestamp from which the CELL_ID assignment became valid. Used to resolve cell site configuration changes in historical joins: filter VALID_FROM_CELL BETWEEN valid_from AND valid_to for point-in-time accuracy.
HOST_NAME
VARCHAR(128)
Destination hostname of the session — SNI (Server Name Indication) or reverse DNS. High cardinality: 500K+ unique domains per day. Dictionary-encoded in Parquet for compression. This is the field that maps to visited_host in the output datasets after domain normalisation (sub-domains collapsed to root domain).
DOWNLOAD_VOLUME
BIGINT (bytes)
Total bytes transferred from server to device (downlink) for this session. Aggregated as SUM at 5-minute and 15-minute windows. Becomes visited_download_data in output datasets. Typical range: 0 to ~50 MB per session.
UPLOAD_VOLUME
BIGINT (bytes)
Total bytes transferred from device to server (uplink). Typically 10–30% of DOWNLOAD_VOLUME for standard browsing; elevated for transactional submissions (bet placement, form posts, account funding). Becomes visited_upload_data in output datasets.
SESSION_DURATION
INT (seconds)
Total duration of the TCP session in seconds, including active transfer and idle keep-alive. Sessions under 5 seconds are typically DNS lookups or pre-fetch traffic. Converted to minutes and aggregated as visited_duration_total in output datasets; divide by 15 to derive session count.
Hit
INT
Number of times the subscriber actively requested the SNI hostname within this session. Value of 1 = single page load or DNS hit; 10+ = active browsing or app polling (social feeds, streaming buffering). 0 = session established but no SNI interaction (background sync). NULL = session predates Hit tracking. Becomes visited_event_count in output datasets.
Aggregation levels: Timestamped (58 GB/day · 21 TB/year, Coldline) → 5-min rollup per TOKEN_ID × CELL_ID × window (1.9 GB/day, Nearline) → 15-min rollup (0.7 GB/day, Standard). The 5-min aggregation adds session_count, distinct_hosts and drops individual HOST_NAME per session — a 31× row reduction.
Source 02
Location Events
Secondary · Daily

Geo-timestamped location pings generated as subscriber devices transition between radio cell coverage areas. Each row is one location event with start and end coordinates. Raw volume is extremely high due to stationary ping repetition (~60–70% of raw pings are identical consecutive coordinates for stationary devices) — deduplication at source reduces ~8 billion raw pings/day to ~2.5 billion unique location events before storage. Uses the same TOKEN_ID hash as Source 01, enabling direct join between browsing behaviour and physical movement without any PII exposure. Parquet with delta-encoded coordinates (consecutive nearby values compressed 4–6×).

Raw pings/day
~8 B
~2.5B after dedup
Raw bytes/row
~88 B
Compressed/row
~15 B
Delta-encoded coords
Daily (compressed)
~37 GB
~13.5 TB/year
FieldTypeDescription
TOKEN_ID
VARCHAR(64)
Privacy-preserving hashed subscriber identifier — identical token format as Source 01. The shared TOKEN_ID is the primary cross-domain join key: joining this table to Source 01 on TOKEN_ID enables direct behavioural + movement enrichment with no PII traversal. Re-salted at midnight UTC; consistent within a calendar day.
start_time
TIMESTAMP TZ
Timestamp when the device entered this geographic position, in UTC. Forms the lower bound of the dwell window. Used for 5-minute and 15-minute bucket assignment in aggregated tables.
end_time
TIMESTAMP TZ
Timestamp when the device left this geographic position. end_time - start_time gives dwell duration. For instantaneous pings (same second), start_time equals end_time — these are network-forced location updates rather than movement events.
start_latitude
DOUBLE (float64)
WGS-84 latitude of device position at start_time, decimal degrees, 15 significant figures (~1mm accuracy). In aggregated tables replaced by H3 grid cell index. High stationarity in source: observed same coordinates 30+ consecutive times for urban dwellers — delta encoding compresses to 2–3 bytes for repeated values.
start_longitude
DOUBLE (float64)
WGS-84 longitude at start_time. Stored separately from latitude for columnar compression efficiency — Parquet delta encoding on consecutive nearby coordinates achieves 4–8× reduction.
end_latitude
DOUBLE (float64)
WGS-84 latitude at end_time. Equals start_latitude for stationary pings. Dropped in 5-minute aggregation in favour of centroid lat/lon or H3 index.
end_longitude
DOUBLE (float64)
WGS-84 longitude at end_time. Dropped in 5-minute aggregation.
cell_site_id
VARCHAR(20)
Radio cell tower serving the device at ping time. Same format as Source 01 CELL_ID — enabling cell-level join across both datasets. Provides network-topology-based geographic anchor when GPS precision is not required. NULL only for Wi-Fi calling or IP-only sessions with no assigned radio cell.
Aggregation levels: Timestamped (37 GB/day, Coldline) → 5-min per device × H3-r8 hexagon (461m cell) → 15-min per device × H3-r7 hexagon (5.16 km²). The 5-min aggregation replaces 4 float64 coordinate columns with one H3 index — saving ~24 bytes/row and making spatial joins trivial. At 15-min aggregation TOKEN_ID may be dropped, enabling fully anonymised k≥5 grouping for BI-safe footfall density reporting.
Source 03
Subscriber Demographics
Reference · Monthly

One row per active subscriber per monthly snapshot. Contains stable identity and socioeconomic attributes derived from billing records, supplemented by third-party income banding. Slowly changing — most fields are constant month-to-month. A Type-2 SCD (Slowly Changing Dimension) approach is used: changed records generate a new row, enabling historical attribute join accuracy. The join key to Sources 01 and 02 passes through an internal linkage layer — the raw subscriber number is never present in the analytical tables.

Rows
~60M
Active subscribers
Raw bytes/row
~81 B
Monthly (compressed)
~2.8 GB
Annual
~34 GB
all monthly snapshots
FieldTypeDescription
Hashed_MSISDN
VARCHAR(64)
Primary join key — SHA-256 hash of the subscriber's mobile number. The linkage between this table and Sources 01/02 is mediated by an internal identity layer: TOKEN_ID in web traffic and location tables maps here through a secure linkage hub. Raw phone numbers are never stored in the analytical layer.
MOBILE_COUNTRY_CODE
SMALLINT
ITU-T E.212 Mobile Country Code identifying the subscriber's home network country. Foreign key to network operator reference. 2 bytes.
MOBILE_NETWORK_CODE
SMALLINT
ITU-T E.212 Mobile Network Code identifying the specific operator within the MCC. Combined MCC+MNC uniquely identifies the PLMN (Public Land Mobile Network). 2 bytes.
Age
TINYINT
Subscriber age stored as a 5-year band integer for privacy compliance (raw date-of-birth not stored): 1=18–24 · 2=25–34 · 3=35–44 · 4=45–54 · 5=55–64 · 6=65+. Updated annually on subscriber anniversary. This field is the source of the age_band dimension in Dataset 02 (output).
Gender
CHAR(1)
Self-reported gender from subscriber profile. Values: M · F · N (non-binary / not specified) · U (unknown). Approximately 3% of records are U. Treated as sensitive — access controlled by role-based permissions. Source of the gender dimension in Dataset 03 (output).
Home_Postcode
VARCHAR(10)
Billing address postcode of the subscriber. Foreign key to the postcode area reference table for area-level attribute enrichment (median income, urban/rural classification, population density). Updated when subscriber changes billing address. The outward code prefix of this field becomes the postcode_area dimension in Dataset 04 (output).
Income
TINYINT
Household income band derived from postcode-level median income (ONS MSOA data) combined with device plan tier: 1=under £20K · 2=£20–35K · 3=£35–50K · 4=£50–75K · 5=£75–100K · 6=over £100K. ~12% of records are NULL (pre-pay subscribers with insufficient signals for income assignment). Refreshed quarterly by data vendor. Source of the income_bucket dimension in Dataset 05 (output).
Join path: To enrich web traffic with subscriber demographics: TOKEN_ID → [linkage hub] → Hashed_MSISDN. The linkage hub is never exposed in the analytical layer — downstream joins use a pre-joined view. SCD Type-2 adds ~5% row growth per month as subscriber attributes change.
Source 04
Reference Master Tables
Reference · Monthly

Static lookup tables providing geographic, network topology, and classification context. Updated monthly when cell site commissioning records change or area demographic data is refreshed. Small in volume but critical — they are the geographic spine connecting all other datasets. Two tables: the cell site master (linking radio cell IDs to physical locations and postcode areas) and the postcode area demographics table (linking postcode prefixes to population, median income, and urban/rural classification).

Cell Site Master  ·  ~500K sites · ~12 MB/mo compressed
FieldTypeDescription
cell_id
VARCHAR(20)
Unique cell identifier. Matches CELL_ID in Source 01 and cell_site_id in Source 02. Format: {site_code}-{year_commissioned}.
site_latitude
DOUBLE
WGS-84 latitude of cell tower antenna. Used to spatially join location pings to sites when GPS is unavailable.
site_longitude
DOUBLE
WGS-84 longitude of cell tower antenna.
postcode_area
VARCHAR(10)
Postcode area of the cell site. Bridge key linking web traffic (via cell) to postcode area demographics without exposing subscriber identity.
valid_from
DATE
Date from which this cell configuration is valid. Enables point-in-time joins for historical analysis.
valid_to
DATE
Date until which this cell configuration is valid. NULL = currently active.
technology
VARCHAR(10)
Radio access technology: 4G-LTE · 5G-NR · 5G-mmWave · 3G-UMTS.
sector
TINYINT
Antenna sector (1–6). Each physical site may have multiple sectors covering different azimuth ranges.
Postcode Area Demographics  ·  ~9K areas · <1 MB/mo
FieldTypeDescription
postcode_area
VARCHAR(10)
Outward code prefix (e.g. SW, M, LS). Primary key. Links to subscriber demographics Home_Postcode prefix and cell site postcode_area.
region_name
VARCHAR(40)
ONS standard region name (e.g. North West, Yorkshire and The Humber, London).
population
INT
Resident population from ONS mid-year estimates.
median_income
INT
Median gross annual household income (GBP) from ONS MSOA-level income estimates.
urban_rural_code
TINYINT
ONS Rural-Urban Classification: 1=Major Urban · 2=Large Urban · 3=Other Urban · 4=Significant Rural · 5=Rural 50 · 6=Rural 80.
centroid_lat
FLOAT
Geographic centroid latitude of the postcode area boundary. Used for radius-based spatial queries.
centroid_lon
FLOAT
Geographic centroid longitude.
Full geographic enrichment path: Source 01 CELL_ID → cell site master cell_id → postcode_area → postcode area demographics. This chain derives area-level demographic context from network topology alone — no subscriber-level location data required. Combined with the subscriber-level join via Hashed_MSISDN, it enables the full demographic segmentation seen in Datasets 02–06.

Data Flow Overview

Feeds arrive via dedicated GCS buckets with transfer encryption. Pub/Sub triggers downstream Dataflow jobs within 15 minutes of file landing. Output datasets are available T+3.

📥
Operator SFTP
Encrypted transfer
›
🪣
GCS Raw Bucket
Parquet · archive tier
›
⚡
Pub/Sub Trigger
<15 min latency
›
🔄
Dataflow ETL
Dedup · normalise · agg
›
🔗
Linkage Hub
TOKEN_ID → Hashed_MSISDN
›
📊
BigQuery
Partitioned · T+3
Cost Estimates · Jul 2026 → Dec 2030

Infrastructure Cost Model

Full operational cost for running the daily BigQuery Notebook ETL pipeline on 16 TB/day of raw telco input (4 TB web + 12 TB foot traffic), materialising all 11 output datasets, and delivering to clients. Raw landing storage excluded — three stages costed: processing, BQ output storage, and data distribution.

Daily Pipeline Input
16 TB
4 TB web · 12 TB foot traffic
Processing Baseline/mo
$17,500
on-demand · 80 TB/day · incl. 5 rerun days
BQ Storage (steady-state)
$146
Capacitor · 12-month window
5-Yr Best Case
$433K
Slots + Storage + S3 Delivery

Annual Cost Breakdown · Enterprise Slots ★ · by Component ($K/yr)

Stacked by cost component · includes 5 rerun days/month · S3 egress at current slider setting · 15%/yr growth

GCP compute CUDs save 25–57%; drag to model discount scenarios
0%1020304050607080%

Stage 1 — BQ Notebook Pipeline · Processing

Daily ETL reads 16 TB of raw telco feeds and materialises all 11 output tables. Each web table (DS01–DS05) scans the full 4 TB web corpus; each foot-traffic table (DS06–DS10) scans the full 12 TB FT corpus. Total effective scan: 5 × 4 TB + 5 × 12 TB = 80 TB/day. Panel (DS11) is a negligible aggregation. On-demand billing is per TB scanned — reserved slots pay a flat hourly rate regardless of scan volume.

Web ETL
DS01 – DS05
5 × 4 TB = 20 TB
$125/day
FT ETL
DS06 – DS10
5 × 12 TB = 60 TB
$375/day
Panel
DS11
~0.1 TB (negligible)
~$0.63/day
Total
11 Outputs Daily
~80 TB total scan/day
$500/day
Pricing ModelRateBaseline/mo 2026 · 6 mo20272028202920305-Yr Total
On-demand
$6.25/TB · no commitment · ~80 TB/day · 5×4 TB web + 5×12 TB FT · incl. 5 rerun days/mo
$500/day$17,500 $105,000$241,500$277,725$319,384$367,291$1,310,900
Enterprise Slots ★
700 slots · 5 hr/day · $0.044/slot-hr · incl. 5 rerun days/mo · scales 15%/yr
$154.00/day$5,390 $32,340$74,382$85,539$98,370$113,126 $403,757
Slots saving vs. on-demand ↓$12,110 $72,660$167,118$192,186$221,014$254,165 $907,143

★ BQ Enterprise Edition at $0.044/slot/hour. 700 slots × 5 hr = $154.00/day — reserved capacity handles all 11 output tables with predictable throughput. Slots cost ($154/day) is 69% less than on-demand ($500/day) — reserved slots pay per slot-hour, not per TB scanned, saving ~$907K over 5 years. 5 additional rerun days/month included (35 compute-days/month = $5,390/month). Slot count scales at 15%/yr. S3 egress costed separately in Stage 3.

Stage 2 — BQ Output Storage · 11 Datasets · 12-Month Rolling Window

Daily output across all 11 tables totals ~80 GB/day uncompressed — web tables contribute ~60 GB, foot traffic ~20 GB. BQ Capacitor auto-compresses at ~3× to 27 GB/day stored. Active storage (≤6 months / 180 days) $0.02/GB/month; long-term (6–12 months) $0.01/GB/month. 12-month partition expiry auto-enforced.

Compression VariantRatioDaily Stored 12-mo VolumeActive /mo (180 d)Long-Term /mo (180 d)Steady-State Total/mo
BQ Capacitor (auto) ★
Native columnar · zero export overhead
~3×27 GB9.9 TB 4,860 GB × $0.020 = $97.20 4,860 GB × $0.010 = $48.60 $145.80
Parquet / Snappy
GCS or S3 export — daily transfer overhead
~5×16 GB5.8 TB S3 Standard: 5,800 GB × $0.023/GB = $133/mo $133 + $41 egress
ZSTD-9
GCS or S3 export · best ratio · slower decode
~8×10 GB3.7 TB S3 Standard: 3,650 GB × $0.023/GB = $84/mo $84 + $26 egress
BQ Capacitor Storage 2026 · 6 mo20272028202920305-Yr Total
Active storage (<6 months / 180 days) $340$940$1,540$1,770$2,040$6,630
Long-term storage (6–12 months) $0$760$700$810$930$3,200
Total BQ Storage $340$1,700$2,240$2,580$2,970$9,830

Storage grows at 15%/yr as subscriber panel and domain/POI coverage expand. Long-term pricing activates automatically after 180 days. Data beyond 12 months is purged via BQ partition expiry — no manual cleanup needed.

Stage 3 — S3 Daily Delivery · Egress + Storage

Output datasets are pushed daily to a client-owned S3 bucket via GCP Storage Transfer Service. Two cost components: GCP → internet egress at $0.085/GB and S3 storage for the 24-month rolling Parquet archive at $0.023/GB/month. Use the slider to model different daily output volumes — from a few GB of filtered signals up to half a TB of full-universe exports.

Daily compressed output to S3
16 GB/day
· 0.48 TB/month · 5.8 TB/year
First month cost
$41/mo
1 GB 512 GB (0.5 TB)
1 GB128 GB256 GB384 GB512 GB
GCP → S3 Egress (Month 1)
$41/mo
$0.085/GB · daily × 30 days
S3 Storage (Month 1, ramps to 24-mo)
$11/mo ↗
$0.023/GB/mo · accumulates over 24 months
S3 Storage at Steady State (24 months)
$269/mo
daily × 730 days × $0.023/GB/mo
5-Year Delivery Total
—
Egress + S3 storage · grows 15%/yr
Alternative — BigQuery Analytics Hub ($0 additional): If clients are GCP-native, they query the linked dataset directly — zero egress and no S3 storage. Publisher pays only for the BQ output storage already counted in Stage 2.

5-Year Cost Summary · Jul 2026 – Dec 2030

Cost Component 2026 · 6 mo20272028202920305-Yr Total
Processing — On-demand $105,000$241,500$277,725$319,384$367,291$1,310,900
Processing — Enterprise Slots $32,340$74,382$85,539$98,370$113,126$403,757
BQ Output Storage (Capacitor) $340$1,700$2,240$2,580$2,970$9,830
Distribution — S3 Daily Push (slider-driven) ——————
On-demand + S3 ——————
Enterprise Slots + S3 Delivery ★ ————— —

All costs grow at 15%/yr. Compute rows include 5 rerun days/month. S3 delivery row is driven by the slider in Stage 3 — adjust there to see 5-year impact. At 700 slots · 5 hr/day ($154/day), Enterprise Slots saves ~$907K vs on-demand ($500/day) over 5 years — paying per slot-hour is significantly more efficient than per-TB billing for this 80 TB/day workload.

Annual Cost by Compute Scenario ($K/yr) — updates with egress slider

On-demand vs Enterprise Slots · includes BQ output storage + S3 delivery at current slider setting

Pricing basis (Jul 2026 GCP/AWS list): BQ on-demand $6.25/TB scanned · BQ Enterprise Slots $0.044/slot/hour · BQ active storage $0.020/GB/month · BQ long-term storage $0.010/GB/month · GCP→internet egress $0.085/GB · AWS S3 Standard $0.023/GB/month. 15%/yr data growth assumed throughout. Committed-use discounts, enterprise negotiated pricing, and pipeline efficiency gains may materially reduce actuals.
Data Outputs

Aggregated Signal Datasets

Five pre-computed BigQuery datasets covering domain-level engagement enriched with funnel signals, and three demographic cuts (age, gender, postcode) plus a panel coverage reference. All datasets are partitioned by date, available T+3, and delivered as read-only shared views — no ETL work required on the client side.

Datasets
11
5 web · 5 foot · 1 reference
Refresh
Daily
T+3 availability
History
24 mo
Rolling lookback window
Delivery
BigQuery
Shared dataset · read-only access
🌐
Section A
Web Traffic
Domain-level engagement signals derived from subscriber clickstream data — enriched with funnel stage classification, Alpha Point flags, and 5 demographic cuts (age, gender, postcode, income, and the base enriched view).
5
Datasets
Dataset 01
Enriched Domain Engagement
Core · Enriched

The primary output dataset. One row per domain per day — aggregated from raw URL-level telco records and enriched with funnel stage classification and Alpha Point flags. This is the table that drives the daily visitor trend chart, Signal Breakdown, and KPI Correlation sections of every Alpha Signal Report. It combines raw engagement metrics (visitors, hits, data volume, duration) with analytical overlays assigned by the signal classification engine.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date on which subscriber activity was observed. All trend and seasonality analysis keys off this field. Filter on visited_date in every query to avoid full-table scans.
visited_host
STRING
Normalised top-level registered domain (e.g. fanduel.com, betfair.com). Sub-domains are collapsed to their root domain during ETL. Primary join key to the ticker-to-domain mapping used in report generation.
signal_stage
STRING
Funnel stage classification assigned to this host on this date by the signal engine. Values: new_user_acquisition · login · open_app · marketing · user_engagement · commerce_intent · transactional_activity · loyalty_activity · churn_risk · unclassified. A single domain may appear under multiple stages on the same date when different URL paths within it are classified differently — aggregate to domain + date to get totals across all stages.
is_alpha_point
BOOLEAN
True when this host+stage combination has been manually verified as a high-confidence predictive signal — one whose historical traffic has demonstrated consistent lead correlation with a reported financial KPI across at least four quarters. Alpha Points are the inputs to the KPI Trend overlay and Signal Correlation sections. Approximately 8% of classified host+stage rows carry this flag.
visitor_total
INTEGER
Count of distinct subscriber tokens that generated at least one network event on this host on this date, within this signal stage. The primary volume signal — equivalent to daily active users (DAU) at domain level. Used as SUM(visitor_total) across dates to derive total active users over a window, or SUM(visitor_total) / COUNT(DISTINCT visited_date) for daily averages.
visited_event_count
INTEGER
Total network events (DNS resolutions, HTTP requests, API calls, asset fetches) attributed to this host and stage on this date. A high visited_event_count / visitor_total ratio indicates interactive, multi-step engagement rather than passive single-page visits. Used in: SUM(visited_event_count) for total hits; divided by visitor_total for hits-per-user.
visited_download_data
INTEGER
Aggregate downstream bytes transferred from the domain to subscriber devices on this date within this stage. Reported in raw bytes — divide by 1,024 for KB or 1,048,576 for MB. Heavy downloads are characteristic of content-streaming and odds-feed stages; informational marketing pages generate low values. Compute KB-per-user as SUM(visited_download_data) / SUM(visitor_total) / 1024.
visited_upload_data
INTEGER
Aggregate upstream bytes from subscriber devices to domain servers. Upload volume is the strongest diagnostic for transactional flows — form submissions, bet placements, payment confirmations, and account-funding operations all produce elevated upload relative to the same domain's other stages. Monitor the visited_upload_data / visited_download_data ratio as a directional indicator of purchase intent.
visited_duration_total
INTEGER
Total subscriber time (minutes) connected to this host and stage across all network events on this date. Includes active transfer periods and idle keep-alive windows. Divide by visitor_total for average minutes per user. Divide by 15 to derive estimated session count using the standard 15-minute inactivity threshold: SUM(visited_duration_total) / 15 = estimated sessions.
Summing visitor_total across all signal_stage values for a given visited_date + visited_host gives the all-stages domain total used in the daily trend chart. The Signal Breakdown table in the report is produced by grouping this dataset by signal_stage.
Dataset 02
Domain Engagement · Age Band
Demographic · Age

The enriched domain engagement metrics from Dataset 01, further broken down by subscriber age band derived from billing records. Age cohort analysis reveals whether new user acquisition is skewing younger or older quarter-on-quarter, which demographic drives transactional volume, and whether churn risk is concentrated in a specific age group. For regulated sectors (sports betting, financial services), the 18–24 cohort split is also material for compliance monitoring.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date of activity. Always filter on this field.
visited_host
STRING
Normalised root domain. Foreign key to Dataset 01.
signal_stage
STRING
Funnel stage classification — same values as Dataset 01.
age_band
STRING
Subscriber age group derived from billing date-of-birth. Stored as a band (not exact age) for privacy compliance. Values: 18_24 · 25_34 · 35_44 · 45_54 · 55_64 · 65_plus · unknown. The 25–44 band consistently generates the highest upload volume on commerce-classified domains, correlating with peak transaction frequency. The 55–64 band shows the highest duration per visitor, suggesting habitual, unhurried engagement.
visitor_total
INTEGER
Distinct subscriber count in this age band accessing the host on this date within this stage. Summing across all age_band values for a given date, host, and stage reproduces the visitor_total from Dataset 01.
visited_event_count
INTEGER
Total network events from this age cohort. Hits-per-user ratio varies substantially by age band — 18–34 users exhibit more frequent, shorter interaction bursts; 45+ users show fewer, longer interactions.
visited_download_data
INTEGER
Downstream bytes for this age band (raw bytes). Download intensity tends to peak in the 25–44 band for streaming and odds-feed domains.
visited_upload_data
INTEGER
Upstream bytes for this age band. The 25–44 cohort generates the highest per-user upload on transactional domains — indicative of highest average stake value.
visited_duration_total
INTEGER
Total connected time (minutes) for this age band. Divide by visitor_total for avg mins per user; divide by 15 for estimated sessions. The 55+ cohort consistently shows the highest duration-per-visitor despite lower raw visitor counts.
Summing across all age_band values for a given visited_date + visited_host + signal_stage triple reproduces the corresponding row in Dataset 01.
Dataset 03
Domain Engagement · Gender
Demographic · Gender

The enriched domain engagement dataset cut by subscriber gender, derived from billing records. Gender-split analysis is particularly material for consumer brands where the product audience composition is a leading indicator of revenue mix — for example, a sustained shift in the male-to-female ratio within the commerce intent stage often precedes changes in reported product revenue for sports betting operators. Also used to validate that product reach is broadening or narrowing across acquisition campaigns.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date of activity.
visited_host
STRING
Normalised root domain. Foreign key to Dataset 01.
signal_stage
STRING
Funnel stage classification — same values as Dataset 01.
gender
STRING
Subscriber gender derived from billing records. Values: M (male), F (female), U (unknown / not provided). The U bucket is non-trivial — typically 10–15% of the subscriber base — and should be excluded from ratio calculations or treated as a separate cohort. For sports betting domains the M share of transactional_activity stage consistently exceeds its share of user_engagement, indicating proportionally higher conversion from engagement to transaction among male subscribers.
visitor_total
INTEGER
Distinct subscribers of this gender accessing the host within this stage on this date.
visited_event_count
INTEGER
Total network events from this gender cohort. Events-per-visitor ratios can differ by 20–30% across gender for the same domain.
visited_download_data
INTEGER
Downstream bytes for this gender cohort (raw bytes).
visited_upload_data
INTEGER
Upstream bytes. Upload volume split by gender is the most direct proxy for transaction count split available without payment data.
visited_duration_total
INTEGER
Total connected time (minutes) for this gender cohort. Divide by 15 for estimated sessions; divide by visitor_total for avg minutes per user.
Summing M + F + U rows for a given date, host, and stage reproduces the Dataset 01 totals. Exclude gender = 'U' when computing M:F ratios to avoid skewing the proportion.
Dataset 04
Domain Engagement · Postcode Area
Geographic

The enriched domain engagement dataset cut by subscriber home postcode area, derived from billing address (not location at time of visit). Geographic segmentation enables spatial concentration analysis — identifying whether consumer activity is nationally distributed or concentrated in specific regions. Regional signals frequently lead national KPI movements: a sustained uplift in transactional upload volume from a specific postcode area often reflects a regional promotional campaign or localised competitor withdrawal before it is visible in aggregate national metrics.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date of activity.
visited_host
STRING
Normalised root domain. Foreign key to Dataset 01.
signal_stage
STRING
Funnel stage classification — same values as Dataset 01.
postcode_area
STRING
Outward code prefix of the subscriber's billing address postcode (e.g. SW, M, LS, B, G). Represents approximately 120 distinct postcode areas covering all of Great Britain and Northern Ireland. This is the subscriber's registered home area — it is static per subscriber and does not change with physical movement. London postcode areas (EC, WC, E, N, NW, SE, SW, W) are retained as separate codes and can be grouped for London-wide analysis. NW England (M, OL, SK) and Yorkshire (LS, BD, HX) consistently over-index for sports betting relative to subscriber base share.
visitor_total
INTEGER
Distinct subscribers registered in this postcode area who accessed the host within this stage on this date.
visited_event_count
INTEGER
Total network events from subscribers in this postcode area.
visited_download_data
INTEGER
Downstream bytes (raw) from this postcode area cohort.
visited_upload_data
INTEGER
Upstream bytes from this postcode area. Regional upload spikes on transactional_activity domains frequently pre-date regional promotional announcements by 5–10 days.
visited_duration_total
INTEGER
Total connected time (minutes) for subscribers in this postcode area. Divide by 15 for estimated sessions; divide by visitor_total for avg minutes per user.
Home postcode area reflects billing address, not physical location at time of visit. Summing all postcode areas for a given date, host, and stage reproduces Dataset 01. For city-level roll-ups, group postcode areas by their standard regional boundaries.
Dataset 05
Domain Engagement · Income Bucket
Demographic · Income

The enriched domain engagement dataset cut by subscriber estimated household income bucket, derived from a propensity model that combines billing plan tier, billing postcode area median income, and device class. Income segmentation enables spend-capacity analysis without transactional data — identifying whether high-value engagement (commerce intent, transactional activity) is concentrated among higher-income cohorts, and how income-sensitive a product's audience acquisition is across quarters. For consumer financial services and premium subscription products, income bucket is the most commercially actionable demographic dimension in the dataset.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date of activity. Always filter on this field.
visited_host
STRING
Normalised root domain. Foreign key to Dataset 01.
signal_stage
STRING
Funnel stage classification — same values as Dataset 01.
income_bucket
STRING
Estimated household income tier derived from a propensity model combining plan tier, billing postcode median income (ONS MSOA-level), and handset class. Values: under_20k · 20k_35k · 35k_50k · 50k_75k · 75k_100k · over_100k · unknown. Buckets reflect gross annual household income in GBP. The unknown bucket covers subscribers where insufficient signals exist to assign a bucket with confidence — typically pre-pay subscribers with no device data. The 50k_75k and over_100k buckets consistently over-index for transactional activity on sports betting and financial services domains relative to their share of the total panel.
visitor_total
INTEGER
Distinct subscribers in this income bucket accessing the host within this stage on this date. Summing across all income_bucket values for a given date, host, and stage reproduces the visitor_total from Dataset 01.
visited_event_count
INTEGER
Total network events from this income cohort. Higher-income buckets tend to generate more events per visitor on commerce and financial domains, reflecting greater product feature utilisation.
visited_download_data
INTEGER
Downstream bytes (raw) from this income cohort. Higher-income subscribers are more likely to use premium or data-heavy product tiers, producing elevated download volumes per visitor.
visited_upload_data
INTEGER
Upstream bytes from this income cohort. Upload volume is the strongest available proxy for transaction frequency and value — the over_100k bucket consistently shows the highest upload-per-visitor ratio on transactional_activity classified domains.
visited_duration_total
INTEGER
Total connected time (minutes) for subscribers in this income bucket. Divide by visitor_total for avg minutes per user; divide by 15 for estimated sessions using the standard telco inactivity threshold.
Summing all income_bucket values for a given visited_date + visited_host + signal_stage reproduces the corresponding row in Dataset 01. Income bucket assignment is a propensity model estimate — treat as directional rather than precise. Use alongside Dataset 11 panel coverage fields to compute income-cohort penetration rates.
📍
Section B
Foot Traffic
Physical location audience intelligence for Points of Interest — daily POI-level footfall with demographic cuts by age, gender, income, and home postcode area. T+4 business day availability.
5
Datasets
Dataset 06
Daily POI Footfall
Core · Physical Location

The primary foot traffic output — one row per Point of Interest per day. Aggregated from raw location ping data, this dataset measures how many distinct subscribers physically visited a given POI on a given day and for how long. Where fewer than 50 unique visitors are observed for a segment, the value is suppressed to −1 to protect individual privacy. This is the base table from which all demographic and geographic foot traffic views are derived. Delivered daily including weekends and public holidays.

FieldTypeDescription
Visited_Point_Of_Interest
STRING
Unique identifier for the Point of Interest. Consistent and non-recycled — the same POI ID will always refer to the same physical location across all historical and future data. Foreign key to the POI Reference table which contains name, address, category, and chain affiliation. Used as the primary dimension for cross-POI comparison and brand-level footfall aggregation.
Visited_Date
DATE
Calendar date on which the visit was observed. Format: YYYYMMDD (e.g. 20250901). Partition key for all foot traffic queries. Daily granularity — the time period covers midnight to midnight local time. Always filter on this field to avoid full-table scans.
Visited_POI_Winner
STRING
Indicates whether this POI has been identified as the primary dwell location for a given subscriber visit, versus a candidate in the top-5 shortlist. Values: Y = confirmed winner POI (subscriber's primary dwell location that day); otherwise the field contains the POI's rank position among up to 5 candidate winners. Used to distinguish confirmed footfall from proximity signals in dense urban environments where a subscriber's device pings may be attributed to multiple adjacent POIs.
Visited_Duration_Total
NUMBER
Total cumulative time (in seconds) that all visitors spent at this POI on this date. Divide by 3,600 for hours; divide by Visitor_Total for average seconds-per-visitor (dwell time). Note: this field is in seconds — unlike the web traffic equivalent which is in minutes. Dwell time is the primary intensity signal for physical location engagement: high dwell with moderate visitor counts indicates a destination POI (e.g. restaurant, gym); low dwell with high visitor counts indicates a transactional or transit POI (e.g. convenience store, transport hub).
Visitor_Total
NUMBER
Total count of distinct subscribers who visited this POI on this date. Each subscriber is counted once regardless of number of visits. Privacy threshold: if the count falls below 50, this field is populated with −1 (N/A) to prevent individual identification. The primary volume signal — directly analogous to daily footfall count. Null-safe aggregations should treat −1 values as suppressed, not as negative counts.
Summing Visitor_Total across all POIs for a given brand (using the chain affiliation from the POI Reference table) produces brand-level daily footfall. Compare against Dataset 01 (Enriched Domain Engagement) on the same date to measure online-to-offline correlation — a rising digital engagement signal ahead of sustained footfall uplift is a leading indicator for physical sales performance.
Dataset 07
POI Footfall · Age Band
Demographic · Age

Daily POI footfall further broken down by visitor age band. Enables cohort-level analysis of which age groups are physically visiting specific locations — critical for retail formats targeting specific demographics, for validating the age composition of a brand's physical audience versus its digital audience (Dataset 02), and for identifying whether demographic mix shifts are occurring at the store or venue level before they aggregate into reported earnings.

FieldTypeDescription
Visited_Point_Of_Interest
STRING
POI identifier. Foreign key to POI Reference table. Primary join key to Dataset 06.
Visited_Date
DATE
Calendar date of visit. Partition key. Format YYYYMMDD.
User_Demographic_Age
NUMBER
Age band of the visiting subscriber cohort, encoded as an integer reference: 1 = 18–24 · 2 = 25–34 · 3 = 35–54 · 4 = 55–75+ · 5 = unknown. Note: the 35–54 band is wider than the equivalent web traffic age bands (35–44 and 45–54) — account for this difference when comparing physical vs. digital audience age composition across datasets.
Visited_Duration_Total
NUMBER
Total dwell time (seconds) for visitors in this age band at this POI on this date. Divide by Visitor_Total for average dwell per visitor. The 25–34 and 35–54 cohorts typically show the longest dwell times for food-and-beverage and leisure POIs; the 18–24 cohort shows the highest visit frequency with shorter average dwell.
Visitor_Total
NUMBER
Distinct visitors in this age band. Suppressed to −1 when fewer than 50 subscribers are in the segment for privacy protection. Summing across all age bands for a given POI and date reproduces the Visitor_Total in Dataset 06 (excluding unknowns).
The age band encoding differs from Dataset 02 (web traffic): band 3 here covers 35–54 combined, versus separate 35–44 and 45–54 bands in the web data. When correlating physical vs. digital age demographics, merge web bands 3 and 4 before comparison.
Dataset 08
POI Footfall · Gender
Demographic · Gender

Daily POI footfall split by visitor gender. Physical gender composition at a store or venue often differs from the brand's digital audience mix — a brand may have a predominantly male online audience (as measured in Dataset 03) but a more balanced or female-leaning physical visitor base. Tracking gender composition of physical visits alongside digital engagement signals enables detection of audience mix shifts at individual POI level before they appear in aggregated trading metrics.

FieldTypeDescription
Visited_Point_Of_Interest
STRING
POI identifier. Foreign key to POI Reference table.
Visited_Date
DATE
Calendar date of visit. Partition key. Format YYYYMMDD.
User_Demographic_Gender
NUMBER
Gender of the visiting subscriber cohort, encoded as a reference integer: 1 = Male · 2 = Female · 3 = Unknown. The unknown bucket typically represents 10–15% of the subscriber panel and should be excluded from M:F ratio calculations. When comparing to Dataset 03 (web gender), the same encoding applies: align on codes 1 and 2 and exclude code 3 from ratio analysis.
Visited_Duration_Total
NUMBER
Total dwell time (seconds) for this gender cohort at this POI on this date.
Visitor_Total
NUMBER
Distinct visitors of this gender. Suppressed to −1 when fewer than 50 in segment. Summing codes 1 + 2 + 3 reproduces Dataset 06's Visitor_Total for the same POI and date.
Dataset 09
POI Footfall · Income Range
Demographic · Income

Daily POI footfall broken down by visitor estimated household income range — a third-party appended attribute derived from billing postcode median income and subscriber plan tier. Income segmentation of physical visits is particularly valuable for premium retail, financial services branches, and hospitality venues where spend capacity directly predicts transaction value. Compare with Dataset 05 (web income) to identify income cohorts that engage digitally but do not convert to physical visits — an indicator of conversion friction or geographic accessibility barriers.

FieldTypeDescription
Visited_Point_Of_Interest
STRING
POI identifier. Foreign key to POI Reference table.
Visited_Date
DATE
Calendar date of visit. Partition key. Format YYYYMMDD.
User_Demographic_Income
NUMBER
Estimated household income band of the visiting subscriber cohort, encoded as a reference integer: 1 = under £25K · 2 = £25K–£74,999 · 3 = £75K–£149,999 · 4 = £150K–£249,999 · 5 = £250K+ · 6 = unknown. Income is a third-party appended attribute — treat as directional rather than precise. The unknown bucket (code 6) covers subscribers without sufficient billing data for income assignment; typically higher among pre-pay subscribers.
Visited_Duration_Total
NUMBER
Total dwell time (seconds) for this income cohort at this POI on this date. Higher-income bands typically show longer dwell times at premium and destination retail formats; lower-income bands show higher frequency with shorter dwell at transactional and convenience formats.
Visitor_Total
NUMBER
Distinct visitors in this income band. Suppressed to −1 when fewer than 50 in segment. Summing across all income codes reproduces Dataset 06's total.
Dataset 10
POI Footfall · Postcode Area
Geographic

Daily POI footfall broken down by visitor home postcode area — where the subscriber lives, not where they are visiting from at the moment of the visit. This distinction matters: a retail store in Manchester will draw visitors from a wide range of home postcode areas; the catchment map produced by this dataset shows the geographic reach of that physical location. Cross-referencing home postcode area with the store's own postcode area quantifies the proportion of local versus out-of-area visitors — a valuable input to site selection, marketing geo-targeting, and competitor proximity analysis.

FieldTypeDescription
Visited_Point_Of_Interest
STRING
POI identifier. Foreign key to POI Reference table.
Visited_Date
DATE
Calendar date of visit. Partition key. Format YYYYMMDD.
User_Demographic_Home_Post_Code_Area
STRING
Home postcode area outward-code prefix of the visiting subscriber cohort (e.g. M, LS, SW). Derived from billing address — static per subscriber, not the location at time of visit. Used to construct visitor catchment maps: which postcode areas contribute the most footfall to a given POI. Cross-reference with the Postcode Area Demographics reference table for population-normalised penetration rates.
Visited_Duration_Total
NUMBER
Total dwell time (seconds) for visitors from this home postcode area at this POI on this date. High duration from distant postcode areas indicates deliberate destination travel rather than convenience visits — a signal of strong brand pull.
Visitor_Total
NUMBER
Distinct visitors from this home postcode area. Suppressed to −1 when fewer than 50 in segment. Summing across all postcode areas for a given POI and date reproduces Dataset 06's total. To compute catchment share: Visitor_Total (postcode X) / Visitor_Total (all postcodes).
Home postcode area represents where the subscriber lives, not their location at time of visit. A visitor counted under postcode area M (Manchester) at a London POI is a Manchester resident who travelled to London — not someone whose device was in Manchester at the time. This enables genuine catchment area analysis rather than proximity analysis. For cross-brand comparisons, normalise by the panel count for each postcode area from Dataset 11 (Panel Coverage & Size).
📊
Section C
Panel Size
Daily reference data on subscriber panel composition and coverage. Essential for normalising visitor metrics across time and computing cohort penetration rates across all web and foot traffic datasets.
1
Dataset
Dataset 11
Panel Coverage & Size
Reference

A daily reference dataset recording the size and demographic composition of the subscriber panel used to compute Datasets 01–05. Panel size is not constant — subscribers join and leave the network daily, and the active panel fluctuates with data collection windows. This dataset is essential for normalising visitor metrics across time: a raw visitor count increase from one month to the next is only meaningful when benchmarked against the panel size in each period. Coverage percentage (active panel ÷ total panel) is the primary data quality indicator for any given date.

FieldTypeDescription
visited_date
DATE
Partition key. Calendar date for which panel size is recorded. Join to any of Datasets 01–04 on visited_date to normalise engagement metrics against panel coverage for that specific day.
panel_size_total
INTEGER
Total number of subscriber tokens in the reference panel for this date — the theoretical maximum number of distinct individuals who could appear in the engagement datasets. This figure reflects the contracted subscriber base included in the data share agreement and is updated daily as subscribers are added or removed.
panel_size_active
INTEGER
Number of subscriber tokens that generated at least one billable network event on this date — the observable panel. Days with low active panel (e.g. Christmas Day, system maintenance windows) produce suppressed visitor counts in Datasets 01–05 and should be flagged in trend analysis. Active panel is typically 60–75% of total panel on weekdays, lower on weekends.
coverage_pct
FLOAT
Active panel as a percentage of total panel: panel_size_active / panel_size_total × 100. The primary data quality flag for a given date. Days where coverage_pct falls below a threshold (typically 55%) indicate incomplete data collection and the corresponding rows in engagement datasets should be treated as unreliable for absolute visitor comparisons. Trend analysis should apply a coverage filter before computing period-over-period changes.
panel_male
INTEGER
Count of male subscribers in the active panel on this date. Combined with Dataset 03 gender splits, this enables penetration rate calculation: visitor_total (gender=M) / panel_male gives the share of male subscribers who visited a given domain on a given day.
panel_female
INTEGER
Count of female subscribers in the active panel on this date. Use alongside Dataset 03 for female penetration rates.
panel_income_under_20k
INTEGER
Active panel count for subscribers assigned to the under £20K estimated household income bucket. Use alongside Dataset 05 (web income) to compute income-cohort penetration rates: visitor_total (income_bucket='under_20k') / panel_income_under_20k.
panel_income_20k_35k
INTEGER
Active panel count for the £20K–£35K income band.
panel_income_35k_50k
INTEGER
Active panel count for the £35K–£50K income band.
panel_income_50k_75k
INTEGER
Active panel count for the £50K–£75K income band.
panel_income_75k_100k
INTEGER
Active panel count for the £75K–£100K income band.
panel_income_over_100k
INTEGER
Active panel count for the over £100K income band. The highest-income segment — use to assess whether premium commerce-intent traffic over- or under-indexes relative to this cohort's share of the panel.
panel_income_unknown
INTEGER
Active panel count for subscribers without a model-assigned income band — typically pre-pay subscribers with insufficient signals (~12% of total panel). Exclude this cohort from income-penetration rate denominators.
panel_age_18_24
INTEGER
Active panel count for the 18–24 age band on this date. Use alongside Dataset 02 to compute age-cohort penetration rates — visitor_total (age_band=18_24) / panel_age_18_24 — which normalise for the fact that some age bands are larger than others in the subscriber base.
panel_age_25_34
INTEGER
Active panel count for the 25–34 band.
panel_age_35_44
INTEGER
Active panel count for the 35–44 band.
panel_age_45_54
INTEGER
Active panel count for the 45–54 band.
panel_age_55_64
INTEGER
Active panel count for the 55–64 band.
panel_age_65_plus
INTEGER
Active panel count for subscribers aged 65 and over.
Always join this dataset when computing penetration rates or period-over-period visitor comparisons. A visitor count that grows 5% while the active panel shrinks 8% represents a real penetration increase of ~14% — which the raw visitor metric alone would obscure. Filter to coverage_pct >= 55 before running trend models. For income-cohort penetration rates, divide the income visitor total from Dataset 05 (web) or Dataset 09 (foot traffic) by the corresponding panel_income_* field for the same date.

Platform Summary

📄 View Full Summary

Plain-English guide to all cost models, assumptions, strategies, and caveats across every product page.