Checking access…

Built on Google Cloud
Omni Insights

Omni Insights
Custom Analytics

Omni Insights is a closed-environment custom analytics platform — bringing Web Traffic, Foot Traffic, and Demographics together inside a secure GCP perimeter, then running inferencing and bespoke analytics entirely within that boundary. No data leaves the project; clients receive curated, controlled outputs via BigQuery dataset sharing.

Closed GCP Environment Data Fusion · 3 Sources Custom Inferencing BQ On-Demand Analytics Controlled Output Delivery
Omni Insights — Custom Analytics
Product Overview

Custom Analytics & Closed-Environment Inferencing

Omni Insights is not a dashboard product — it is a controlled analytics environment. Three datasets (Web Traffic, Foot Traffic, Demographics) are brought together inside a VPC-SC GCP perimeter, fused at the subscriber level, and then subject to any combination of custom SQL analytics, ML inferencing, and audience modelling. Results are delivered as curated BigQuery datasets — no raw data ever leaves the project boundary.

🔒
Closed-Loop by Design

All three source datasets — Web Traffic, Foot Traffic, and Demographics — are ingested into and remain within a single GCP project. Inferencing models run inside BigQuery ML or Vertex AI within the same perimeter. Client teams never access raw feeds; they receive only the controlled analytical outputs published to a shared BigQuery dataset. This architecture satisfies data residency requirements and eliminates third-party egress risk.

VPC-SC Perimeter No Raw Data Egress BQ Dataset Sharing Pseudonymised Throughout k-Anonymity Enforced
Core Capabilities — 6 Analytics Modules
🔗
Data Fusion Engine
Joins Web Traffic, Foot Traffic, and Demographics on a shared pseudonymised subscriber key inside BigQuery. Produces a unified subscriber-level profile covering online intent, physical behaviour, and demographic context — the foundation for all downstream analytics.
BQ Join · subscriber_id
🎯
Custom Segmentation
Build bespoke audience segments using any combination of web category interest, physical POI visit patterns, dwell time, age band, and tenure. Segments are defined by client analysts via SQL or a no-code interface and materialised as BQ tables within the closed environment.
BQ ML · Custom Rules
🛒
Attribution Analytics
Measures the digital-to-physical conversion funnel: which subscribers browsed a retailer domain then visited a physical store within a configurable time window. Configurable attribution windows (1–30 days), funnel step analysis, and competitor visit overlap scoring.
Funnel · Time-window
🤖
ML Inferencing Pipeline
Propensity models, churn predictors, and next-best-action classifiers trained and served entirely within the GCP project using BigQuery ML or Vertex AI. Models are retrained on the fused dataset on a scheduled cadence; predictions are written back to BQ as enriched subscriber attributes.
BQ ML · Vertex AI
📊
Ad-Hoc Query Engine
On-demand BigQuery analytics against any combination of the three source datasets for any rolling data window up to 6 months. Partition pruning on date_partition reduces effective scan by ~65%, making large cross-dataset joins cost-efficient. Full SQL flexibility — no schema lock-in.
BQ On-Demand · SQL
📤
Controlled Output Delivery
Analytical results are published exclusively to a client-owned BigQuery dataset via Analytics Hub dataset sharing. Clients query their curated output tables using their own BQ credentials — they never touch raw feeds. Output schemas are agreed upfront and versioned; delivery SLA is configurable per engagement.
Analytics Hub · Sharing
Closed-Environment Data Flow
Data Sources (ingested into perimeter)
Web Traffic
4 TB/day
~50M browsing sessions
Foot Traffic
10 TB/day
~2B location pings
Demographics
~2 TB/month
Monthly subscriber snapshot
Processing — Inside VPC-SC Perimeter
🔗
Data Fusion
BQ join on subscriber_id
›
🎯
Segment & Score
Custom rules + BQ ML
›
🤖
Inferencing
Vertex AI / BQ ML
›
📊
Ad-Hoc Analytics
BQ On-Demand SQL
Controlled Output (no raw data leaves)
Curated BQ Dataset
Analytical outputs published to client's shared dataset via Analytics Hub — pre-aggregated, schema-controlled
Audience Segments
Named segment tables exported in agreed format — subscriber counts only, never individual rows
Model Scores
Propensity and classification scores appended as enriched BQ columns — not individual predictions
Example Use Cases
🛒
Online → In-Store Attribution
Identify subscribers who browsed a retailer domain then visited a physical location — configurable attribution window from 1 to 30 days.
📉
Churn Propensity Scoring
BQ ML classifier trained on tenure, segment, browsing activity drop-off, and reduced physical mobility to score churn risk across the full subscriber base.
🎯
Precision Audience Segments
Combine web category affinity, POI visit patterns, and demographic bands to build high-precision audiences for campaign targeting — all within the closed environment.
📍
Competitor Footfall Overlap
Identify subscribers who visited a target brand location and also visited a competitor — revealing switching behaviour and market share dynamics.
⏱️
Engagement Depth Scoring
Combine web session duration with physical dwell time to produce a composite engagement score per subscriber, per brand or category.
📅
Temporal Trend Analysis
Track how online and offline behaviours shift week-on-week across a rolling 6-month window to surface seasonal patterns and campaign lift.
Data Requirement

Input Data Sources

Two high-velocity daily datasets are joined to produce cross-channel insights. Web Traffic captures online intent; Footfall captures physical presence. Linked by subscriber identifier, together they reveal the consumer journey end-to-end.

🌐
Web Traffic
Daily ingestion · raw CDR browsing sessions
4 TB/day
Raw daily
4 TB
~50M sessions
Monthly raw
120 TB
~1.5B sessions
Monthly compressed
34.3 TB
Parquet · 3.5×
Bytes / row
~80 B
Raw uncompressed
FieldTypeDescription
subscriber_idSTRINGPrivacy-preserving pseudonymised subscriber identifier (SHA-256 of MSISDN + daily salt). Re-salted at midnight UTC — consistent within a calendar day, different across days. Primary cross-domain join key linking this table to Foot Traffic and Demographics.
session_tsTIMESTAMPSession start timestamp in UTC. Precision to the second. Partition basis — date_partition is extracted from this field. Used for 5-minute and 15-minute time-window bucketing in downstream aggregations.
domainSTRINGRoot-level destination hostname of the session derived from SNI (Server Name Indication) or reverse DNS. Sub-domains are collapsed to root during ingestion (e.g. news.bbc.co.uk → bbc.co.uk). High cardinality: 500K+ unique domains per day; dictionary-encoded in Parquet.
categorySTRINGIAB Content Taxonomy 2.0 classification of the domain (e.g. ECOMMERCE, NEWS, SPORTS, FINANCE). Mapped from a 500K-domain lookup table refreshed monthly. NULL for uncategorised or newly registered domains (~4% of sessions).
duration_secINT64Total TCP session duration in seconds including active transfer and idle keep-alive. Sessions under 5 seconds are typically DNS lookups or pre-fetch background traffic. Converted to minutes and aggregated as duration_total in downstream output tables.
bytes_downINT64Downstream bytes transferred from server to device for this session. Aggregated as SUM at 5-minute and 15-minute windows. Typical range: 0 to ~50 MB per session. Streaming and large-file downloads produce outliers; winsorised at 99th percentile in analytical outputs.
bytes_upINT64Upstream bytes transferred from device to server. Typically 10–30% of bytes_down for standard browsing. Elevated for transactional sessions such as form submissions, file uploads, video calls, and bet placements — making it a useful signal for high-intent commercial interactions.
cell_idSTRINGUnique identifier of the serving radio cell at session start. Format: {site_code}-{year_commissioned}. Foreign key to the cell site master table, enabling network-topology geographic resolution (postcode area, lat/lon of antenna) without storing subscriber GPS directly.
event_countINT64Number of distinct URL requests (SNI hits) made to the domain within this session. 0 = background sync with no user interaction; 1 = single page load; 10+ = active browsing or app polling (social feeds, streaming buffers). NULL for sessions predating event-count instrumentation.
device_typeSTRINGSubscriber handset class derived from network signalling and User-Agent: mobile · tablet · fixed · iot. ~2% of sessions return unknown where UA is absent or obfuscated.
date_partitionDATEPartition key extracted from session_ts. Used for BigQuery partition pruning and rolling-window expiry. Historical data beyond the configured retention window is automatically dropped as new partitions arrive.
Aggregation levels: Raw sessions (4 TB/day, Coldline archive) → 5-min rollup per subscriber_id × cell_id × window, adding session_count and distinct_domains (drops individual domain rows — 30× row reduction) → 15-min rollup per cell_id only, enabling k≥5 anonymised BI reporting. Aggregated tables reside in Nearline and Standard tiers respectively.
📍
Foot Traffic
Daily ingestion · raw GPS / location pings
10 TB/day
Raw daily
10 TB
~200M pings
Monthly raw
300 TB
~6B pings
Monthly compressed
85.7 TB
Parquet · 3.5×
Bytes / row
~50 B
Delta-encoded coords
FieldTypeDescription
subscriber_idSTRINGPrivacy-preserving pseudonymised identifier — identical token format and daily salt rotation as Web Traffic. The shared subscriber_id is the primary cross-domain join key: joining this table to Web Traffic on subscriber_id enables behavioural + movement enrichment with no PII traversal.
ping_tsTIMESTAMPTimestamp when the device entered this geographic position (UTC). Forms the lower bound of the dwell window. Used for 5-minute and 15-minute bucket assignment in aggregated tables and for POI dwell-time calculation against end_ts.
end_tsTIMESTAMPTimestamp when the device left this geographic position. Together with ping_ts, forms the precise dwell window: end_ts − ping_ts gives actual time-at-location in seconds. For instantaneous network-forced pings, end_ts equals ping_ts. NULL for the most recent open ping (device still present).
latitudeFLOAT64WGS-84 latitude of device position at ping_ts, decimal degrees to 6 significant figures (~1 m accuracy at the equator). High stationarity in source: urban subscribers often produce identical coordinates for 30+ consecutive pings. Delta-encoded in Parquet, achieving 4–8× compression for repeated values.
longitudeFLOAT64WGS-84 longitude at ping_ts. Stored as a separate column from latitude for columnar compression efficiency — Parquet delta encoding on consecutive nearby coordinates achieves greater reduction than interleaved storage. Replaced by H3 index at the 5-minute aggregation level.
cell_site_idSTRINGRadio cell serving the device at ping time. Same format as Web Traffic cell_id — enabling direct cell-level join across both datasets. Provides a network-topology geographic anchor when GPS precision is unavailable. NULL only for Wi-Fi calling or IP-only sessions with no assigned radio cell.
poi_idSTRINGMatched point-of-interest ID from the 1.2M-POI master dataset. Assigned during ingestion via H3 spatial join: pings whose lat/lon falls within a POI boundary polygon are tagged at source. NULL for pings that fall outside any POI boundary (open street, residential area, transit corridor).
poi_categorySTRINGPOI classification: retail · transport · leisure · health · hospitality · office. Derived from the POI master table; NULL where poi_id is NULL. Used as the primary dimension for footfall segmentation and cross-channel attribution queries.
dwell_minutesINT64Minutes spent at the matched POI, derived as FLOOR((end_ts − ping_ts) / 60). 0 = transit ping (device passed through without dwelling); NULL = no POI match. Winsorised at 480 minutes (8 hours) to suppress overnight device misclassification.
date_partitionDATEPartition key extracted from ping_ts. Used for BigQuery partition pruning and rolling-window expiry. Historical data beyond the configured retention window is automatically dropped as new partitions arrive.
Aggregation levels: Raw pings (10 TB/day, Coldline archive) → 5-min per subscriber_id × H3-r8 hexagon (461 m cell), replacing 2 float64 coordinate columns with one H3 index and saving ~16 bytes/row → 15-min per H3-r7 hexagon (5.16 km²), with subscriber_id optionally dropped to enable k≥5 anonymised footfall density reporting. Deduplication at source reduces stationary consecutive pings by ~65% before storage.
👥
Demographics
Monthly snapshot · subscriber attributes
~2 TB/mo
Monthly volume
~2 TB
Full subscriber base
Refresh cadence
Monthly
Snapshot + deltas
Privacy
k-anon
Pseudonymised
FieldTypeDescription
subscriber_idSTRINGJoin key — same pseudonymised identity as web & footfall
age_bandSTRINGAge cohort bracket (e.g. 18–24, 25–34, 35–44 …)
genderSTRINGInferred or declared gender classification
segmentSTRINGBehavioural segment label (e.g. urban commuter, suburban family)
nationalitySTRINGISO country code of subscriber nationality
tenure_bandSTRINGLength of relationship: <1yr · 1–3yr · 3–5yr · 5yr+
home_regionSTRINGInferred home region from overnight dwell pattern
snapshot_monthDATEPartition key — month the snapshot was generated
Dataset Join & Output Signals

Both datasets share subscriber_id as the linkage key. Each analysis scans a configurable rolling window (1–6 months) of Parquet-compressed data in BigQuery, applying partition pruning on date_partition. The join produces subscriber-level cross-channel profiles combining online category interest with physical POI visits.

Web Traffic
34.3 TB/mo
compressed Parquet
⊕
Foot Traffic
85.7 TB/mo
compressed Parquet
→
Joined Insight
subscriber_id
cross-channel profile
🛒
Online → In-store Attribution
Identify subscribers who browsed a retailer's site and subsequently visited a physical store within a configurable time window.
📊
Cross-channel Audience Profiles
Build rich audience segments combining web category preferences with physical POI visit patterns for precision targeting.
🔁
Conversion Funnel Analysis
Measure digital-to-physical conversion rates: how many online intent signals resulted in a verified physical visit.
📍
Catchment & Competitor Visits
Analyse where else subscribers visit before and after a target POI — uncovering competitor footfall and travel patterns.
⏱️
Dwell & Engagement Scoring
Combine web session duration with physical dwell time to score subscriber engagement depth for a brand or location.
📅
Temporal Trend Analysis
Track how online and offline behaviours shift week-on-week across a 6-month window to identify seasonal and campaign effects.
Data retention: Raw data is maintained in BigQuery active storage on a rolling 6-month window. Data beyond 6 months transitions to long-term storage at $0.01/GB/month. All fields are pseudonymised — no personally identifiable information is stored or processed.
Cost Estimates

GCP Cost Forecast · Jul 2026 – Dec 2030

Cost is purely BigQuery On-Demand compute — you are querying data that already lives in the ingestion pipeline. No separate storage is charged. Each analysis scans Web + Foot Traffic for the chosen data window; cost scales with how much data you look back over and how often you run.

⚡

Recommended: BigQuery On-Demand

For ad-hoc cross-channel analytics, BigQuery On-Demand (pay-per-TB-scanned) is the right choice. The raw data already exists in the ingestion pipeline — no duplicate storage cost. Zero idle slot cost between runs. BQ's columnar Parquet format with partition pruning on date_partition reduces effective scan by ~65%, keeping large cross-dataset joins cost-efficient.

No idle slot cost ~65% scan reduction via column pruning Partition-aware joins $6.25/TB scanned Scales to any frequency
Applies to BQ On-Demand query cost
0%1020304050607080%

Analysis Parameters

How far back each analysis looks into the existing pipeline data. More months = larger BQ scan = higher query cost. No extra storage is charged — the data already exists.
Each analysis runs a full join of Web + Foot Traffic for the selected data window. Cost scales linearly with frequency.
Final joined & aggregated dataset written to BigQuery after each Python/SQL analysis run. Stored at $0.02/GB/month active storage; older results age to $0.01/GB/month after 90 days.
Per Analysis
—
BQ query compute
Output Storage / mo
—
final datasets in BQ
Annual Total
—
compute + storage
5-Year Total
—
Jul 2026 – Dec 2030

Annual Cost by Year

BQ On-Demand query compute (teal) + output dataset storage (purple) · input data carries no storage cost

Query Compute Output Storage

Year-by-Year Forecast

Cost is entirely query-driven — scales with data window size and analysis frequency. 2026 covers Jul–Dec only (6 months).

Year Months Analyses Query Compute Output Storage Total
Total (5yr) 54 — — — —
Pricing basis: BQ On-Demand $6.25/TB scanned · 35% effective column scan with partition pruning on date_partition · 15% join shuffle overhead included · input data has no storage charge (already in ingestion pipeline) · output dataset: $0.02/GB/month active, $0.01/GB/month after 90 days. Costs exclude BQ result egress (typically <$50/year).
Cost Per Analysis
—
BQ On-Demand query
Data window 3 months
Effective scan —
Analyses/yr 12
Annual compute —
Output storage/mo —
5-yr total —

Platform Summary

📄 View Full Summary

Plain-English guide to all cost models, assumptions, strategies, and caveats across every product page.