Verifying access…

Eureka AI
GCP Cost Intelligence
Eureka AI Β· Platform Overview

Platform Summary

A plain-English guide to every page on this site β€” what each model covers, what the numbers mean, what assumptions were made, and what to watch out for. Use the arrow on the left to jump between sections.

🏠 Home β€” Cost Intelligence Hub

The master dashboard β€” all costs in one place

The home page is your starting point. It combines every product's cost into a single 5-year view so you can see the total GCP bill at a glance β€” and drag a discount slider to model what a negotiated enterprise deal would cost.

5-Year GCP Total
~$2.16M
List price, all products
2026 H2 Cost
$145.8K
Jul – Dec 2026 only
Products Modelled
5
Alpha Β· Enterprise Β· Pinpoint Β· Insights Β· Raw Storage
Horizon
5 yrs
Jul 2026 – Dec 2030
πŸ“–What this page shows

Think of the home page like a company's P&L summary β€” it doesn't explain every line in detail, but it tells you the bottom line quickly. You see four product cost models (Omni Alpha, Omni Enterprise, Omni Pinpoint, Omni Insights) plus raw data storage costs rolled up into one table and two charts.

The 5-Year Technology Costs table breaks costs down by year and by cost line (compute, storage, delivery). The two charts visualise the same data from different angles β€” one grouped by product, one grouped by cost type (storage vs compute vs delivery).

The GCP Committed Use Discount slider lets you drag from 0% to 80% to instantly see how a negotiated deal changes every number. At 0% you see full GCP list prices; at 40% (a common enterprise CUD level) costs drop proportionally across the board.

Strategy: All costs are expressed at GCP list price first, then the discount slider applies a uniform multiplier. This keeps the base model honest β€” you can always see what Google would charge without any deal, and separately model what a negotiated rate saves. 2026 is modelled as H2 only (July–December, 184 days) since that is when the platform goes live.
πŸ“Key assumptions
  • All figures are GCP infrastructure only β€” team salaries, software licences, cloud support contracts, and professional services are excluded.
  • 20% year-on-year data volume growth is assumed for Omni Enterprise (category pipeline handles more subscribers each year). Other products use their own growth assumptions.
  • Intermediate data storage uses a 12-month rolling window (default) for Omni Enterprise β€” configurable from 1 to 12 months on the Enterprise cost page. Shorter windows reduce storage cost; 12 months is the default to support reprocessing and backfill.
  • 2026 costs cover H2 only (6 months) because the platform launches mid-year. All subsequent years are full 12-month periods.
Caveats: The home page totals are static β€” they reflect the default slider positions on each product page (20 categories, 50% coverage, 12-month retention for Enterprise; 12 analyses/year for Omni Insights). If you change sliders on individual product pages, the home page does not automatically update. Think of it as the "agreed baseline" view.

⚑ Omni Alpha

The core analytics engine β€” always on, always running

Omni Alpha is the foundational data product. It processes raw network signals daily and produces structured, queryable outputs. Think of it as the engine room β€” it runs every day whether or not anyone is actively querying it, so costs are relatively predictable.

2026 H2 Cost
$33.1K
Jul – Dec 2026
5-Yr Total
~$433K
2026 H2 β†’ 2030
Compute Model
BQ Slots
700 Enterprise slots Β· 5 hr/day
Distribution
S3 Daily
$0.085/GB GCP egress
πŸ“–What this page shows

Omni Alpha uses BigQuery Enterprise Slots β€” a reserved block of compute capacity you pay for by the hour regardless of how much data you scan. This is different from "on-demand" pricing where you pay per terabyte scanned. Slots are the right choice here because the daily pipeline scans 80 TB (5Γ—4 TB web + 5Γ—12 TB foot traffic across 11 output tables) β€” at $6.25/TB on-demand that's $500/day, whereas 700 reserved slots Γ— 5 hr = $154/day.

The pipeline runs daily and produces all 11 output tables. Outputs are delivered to clients as daily S3 file drops β€” GCP-to-internet egress at $0.085/GB applies.

The cost page shows a full 5-year model with both on-demand and slots scenarios side by side, plus sliders for GCP discount and S3 daily volume.

Strategy: BQ Enterprise Slots decisively beat BQ On-Demand for Omni Alpha at this scan volume. The pipeline scans 80 TB/day across 11 output tables (5Γ—4 TB web + 5Γ—12 TB FT) β€” on-demand billing at $6.25/TB comes to $500/day ($17,500/month). A 700-slot reservation running 5 hr/day costs $154/day ($5,390/month) β€” 69% cheaper, saving ~$907K over 5 years. Outputs are delivered via daily S3 file drops; GCP egress ($0.085/GB) is the only distribution cost.
πŸ“Key assumptions
  • 700 BQ Enterprise Slots at $0.044/slot-hour, running 5 hours per day = $154/day fixed compute. Slot count scales 15%/yr to accommodate 15%/yr data growth.
  • 11 output tables scanning 80 TB/day total β€” 5 web tables each scanning 4 TB, 5 foot-traffic tables each scanning 12 TB, 1 panel aggregation (negligible). On-demand equivalent: $500/day.
  • Output storage retained for 12 months (rolling window). Active tier (≀6 months) at $0.020/GB/month; long-term tier (6–12 months) at $0.010/GB/month.
  • S3 daily delivery at $0.085/GB GCP egress. Delivery volume defaults to 16 GB/day compressed output; adjustable via slider on the cost page.
  • 5 rerun days/month included (35 compute-days/month) to account for pipeline reruns and backfills.
Caveats: Slot costs scale linearly with hours used β€” if the daily 5-hour window is insufficient, adding hours increases cost proportionally ($154/day at 5 hr β†’ $185/day at 6 hr). The 80 TB/day scan figure assumes each table reads its full source partition; partition pruning or clustering improvements could reduce this. S3 delivery cost is driven by daily output volume β€” the 16 GB/day default is a baseline; heavier subscriber slices increase this. Ad-hoc or exploratory queries by data scientists run on BQ On-Demand separately from the slot reservation.

⚑ Omni Enterprise

Category intelligence β€” who uses what, and how much it costs to find out

Omni Enterprise answers: "Which subscribers use Ride Hailing apps? Which use Food Delivery? How does each category grow?" It does this by running a daily pipeline that identifies every subscriber who touched a category's URLs, then pulls their full browsing and location history. The result is category-level intelligence grounded in real network observations β€” not surveys.

2026 H2 Cost
$40.5K
Jul – Dec 2026 Β· per-category mode
5-Yr Total
~$561K
2026 H2 β†’ 2030 Β· BQ Slots
Daily Pipeline
~$220
BQ Slots Β· 500 slots Γ— 10hr
Intermediate Storage
~$40
Per month per category (5 tables Γ— 12 mo)
πŸ“–What this page shows

The Omni Enterprise pipeline has 5 steps that run every day in sequence. First, it looks up which web URLs belong to each category (e.g. Uber, Bolt, Free Now β†’ Ride Hailing). Second, it scans the entire web traffic log to find every subscriber who visited those URLs β€” this produces a list of subscriber IDs (MSISDNs). Third, it pulls the full web history for those subscribers. Fourth, it pulls their full location (foot traffic) history. Fifth, it aggregates everything into category-level metrics.

The pipeline runs in per-category mode by default β€” each of the 20 categories executes independently, scanning 50% of subscribers (41M each). Total daily scan: 20 MSISDN lookups (20 TB) + 20 web reads (40 TB) + 20 foot reads (120 TB) + aggregation (16 TB) = ~196 TB/day. BQ Enterprise Slots (500 slots Γ— 10hr = $220/day) is the most cost-effective option; BQ On-Demand would cost $1,225/day β€” 5.6Γ— more expensive.

The cost page has interactive sliders: number of categories (default 20), average subscriber coverage per category (default 50%), processing mode (per-category default vs batched), compute option (BQ Slots / Dataproc / BQ On-Demand), and the intermediate data retention window.

Strategy: BQ Enterprise Slots is the recommended compute at $220/day (500 slots Γ— 10hr) β€” 5.6Γ— cheaper than BQ On-Demand ($1,225/day, 196 TB Γ— $6.25/TB). Each of 20 categories runs independently in per-category mode, scanning 50% of subscribers β€” total 196 TB/day. Dataproc Serverless costs ~$236/day, making the platforms near-equivalent on cost ($16/day gap). BQ Slots is preferred for its pure-SQL simplicity; Dataproc is a valid alternative for Spark-native teams. 5-yr saving vs Dataproc: ~$40K (BQ Slots ~$561K vs Dataproc ~$601K).
πŸ“Key assumptions
  • 20 categories, 50% average subscriber coverage per category (per-category mode) β€” each category independently scans 41M subscribers (50% of 82M). Total scan across all 20 categories: ~196 TB/day (20 TB MSISDN + 40 TB web + 120 TB foot + 16 TB aggregation).
  • Web partition: 4 TB/day per category (logical scan). BigQuery HOST_NAME clustering means the MSISDN lookup scan only touches ~25% of the partition (~1 TB per category).
  • Foot traffic partition: 12 TB/day per category (at 50% coverage). In per-category mode with 20 categories: 20 Γ— 12 TB Γ— 50% = 120 TB foot scanned per day total.
  • 12-month intermediate data retention, 5 tables per signal type per category β€” Steps 3 & 4 each persist 5 day-level aggregated tables (1 row/subscriber/day) at 5Γ— BQ Capacitor compression. At 41M subscribers per category: ~4.1 GB/day (web) + ~3.3 GB/day (foot) = ~$40/month per category. Retention window adjustable via slider (1–12 months).
  • BQ Enterprise Slots: 500 slots Γ— 10 hours/day at $0.044/slot-hour = $220/day fixed compute cost (per-category mode).
  • 20% year-on-year data volume growth applied to both compute and storage.
Caveats: In per-category mode, compute costs scale with the number of categories (N Γ— 196 TB total at 20 cats). Batched mode can reduce compute costs to ~$18/day but requires a single shared scan. BQ Enterprise Slots compute stays fixed at $220/day in per-category mode regardless of data volume β€” this is the key advantage over BQ On-Demand ($1,225/day), which scales with TB read. At $220/day, BQ Slots is only $16/day cheaper than Dataproc ($236/day), making the two platforms near-equivalent; the choice is primarily about SQL vs Spark team preference. Intermediate storage is day-level aggregated rows (~$40/month per category at 50% coverage). The model assumes BQ Enterprise Slots pricing as of July 2026. Costs do not include raw source data storage (web and foot partitions), modelled separately in the Raw Storage Forecast page.

πŸ“ Omni Pinpoint

POI-level location intelligence β€” 1.2M points, 70M subscribers, daily aggregation

Omni Pinpoint is a standalone BigQuery Flex Slots pipeline that processes 10 TB/day of raw foot traffic from 70M subscribers (~35B rows/day) and aggregates it to POI Γ— hour Γ— demographic and POI Γ— day Γ— demographic output tables for 1.2M Points of Interest. The output is a rich, privacy-safe crowd intelligence layer usable for retail analytics, urban planning, venue attribution, and origin-destination flows.

Daily Compute
$120
500 slots Β· 6 hrs Β· $0.04/slot/hr
Annual Compute
$43,800
2026 baseline Β· 18% YoY growth
Agg Storage / yr
~$224
~969 GB SS Β· Hourly 3-mo + Daily 12-mo rolling
5-Yr Total
~$293K
Compute-dominated Β· storage <1%
πŸ“–What this page shows

The cost estimate page covers three stages. Stage 1 explains the platform choice (BigQuery Flex Slots over On-Demand or Dataproc) and shows the 5-step daily pipeline: raw scan β†’ H3 spatial join β†’ demographic enrichment β†’ hourly aggregation β†’ daily rollup. Stage 2 is an interactive compute calculator β€” adjust slots, processing window, and growth rate to model different configurations. Stage 3 covers aggregated output storage, with sliders for POI count, demographic dimensions, and hourly retention window.

A GCP Committed Use Discount slider sits at the top of the page (0–40%) and applies to compute cost across all sections β€” the same pattern as the other product pages.

Platform rationale: BigQuery Flex Slots ($0.04/slot/hr, 1-minute minimum commitment) is optimal for this workload. On-Demand BQ pricing at $6.25/TB would cost ~$62/day just for the raw 10 TB scan β€” before any joins or aggregations. Dataproc Serverless would add operational overhead for a workload that is purely SQL + GEO. Flex Slots burst during the 6-hour daily window and cost $0 outside of it. At 500 slots Γ— 6 hrs, total slot-hours per day = 3,000 at $0.04 each = $120/day. 500 slots is the practical minimum for reliable 6-hour completion of 35B rows given the H3 coordinate conversion (compute-heavy UDF on every ping) and the 70M-subscriber demographic shuffle-join; 200 slots would require 12–15 hours.
πŸ“Key assumptions
  • 70M subscribers, each generating ~500 location pings/day = ~35B rows/day = 10 TB/day raw foot traffic.
  • 1.2M POIs resolved via the 200 MB Master Dataset (broadcast join β€” no shuffle required).
  • 500 BQ Flex Slots Γ— 6 hrs/day β€” minimum reliable allocation for 35B rows: H3 coordinate conversion (UDF on every ping) + 70M-subscriber demographic shuffle-join complete within the 6-hour window. Cost = 500 Γ— 6 Γ— $0.04 = $120/day. At 200 slots the pipeline requires 12–15 hours β€” incompatible with a daily batch cadence.
  • 16 demographic dimensions per POI-hour cell (e.g. 4 age bands Γ— 2 genders Γ— 2 nationality groups).
  • Hourly table: 1.2M POIs Γ— 24 hrs Γ— 16 dims = 461M rows/day β†’ ~9.2 GB/day Parquet-compressed. 3-month rolling retention = ~829 GB steady-state.
  • Daily table: 1.2M POIs Γ— 16 dims = 19.2M rows/day β†’ ~384 MB/day. 12-month rolling = ~140 GB steady-state.
  • BQ storage tiers: hourly table (3-month rolling, all within active window) billed at $0.020/GB/mo; daily table (12-month rolling) billed as 6 months Active ($0.020) + 6 months Dormant ($0.010). Total ~969 GB steady-state = ~$224/year β€” under 0.6% of annual compute cost.
  • 18% YoY growth in subscriber base and data volume β€” drives both compute slot-hours and storage steady-state over the 5-year forecast.
  • sub_hash is the pseudonymised join key linking foot traffic to the 4M-row demographic table β€” a shuffle join (demographic table is 3.5 GB, too large to broadcast).
πŸ“Š5-Year cost summary
  • 2026 H2: $22,080 compute + $112 storage = $22.2K
  • 2027: $51,684 compute + $265 storage = $51.9K
  • 2028: $60,987 compute + $312 storage = $61.3K
  • 2029: $71,965 compute + $368 storage = $72.3K
  • 2030: $84,918 compute + $435 storage = $85.4K
  • 5-Year Total: ~$293.1K β€” storage contributes $1,492 (0.5% of total).

All years assume 18% YoY data growth and the default 3-month hourly retention / 12-month daily retention window. Use the interactive sliders on the cost page to model alternative configurations.

Caveats: The 500-slot / 6-hour window is the minimum reliable estimate for 35B rows of H3 spatial processing and 70M-subscriber demographic joins. Actual wall time depends on BQ scheduler behaviour and query plan efficiency β€” if the demographic shuffle-join (35B rows Γ— 3.5 GB sub table) is slower than expected, scaling to 750–1,000 slots reduces wall time while keeping daily cost linear ($0.04 Γ— 1,000 Γ— 6 hrs = $240/day at 1,000 slots). The 16 demographic dimensions default is configurable β€” increasing to 32 doubles both hourly and daily aggregation output sizes. Committed Use Discounts for BQ Flex Slot reservations of 20% (annual) to 40% (3-year) apply to compute only.

πŸ” Omni Insights β€” Custom Analytics

Pay-per-analysis β€” run a custom query, pay only for what you scan

Omni Insights is for answering specific business questions that aren't covered by the standard product dashboards. For example: "Which subscribers visited a competitor's store within 7 days of browsing our app?" Each question requires joining months of web and foot traffic data on-the-fly. BigQuery On-Demand billing means you pay only when you run an analysis β€” there's no idle cost.

Per Analysis
~$906
At 3-month window, 12/yr
Annual Compute
~$10.9K
12 analyses Γ— $906
Output Storage
~$0.1K
Per year Β· $0.5K 5-yr total
Input Storage
$0
Data already in pipeline
πŸ“–What this page shows

Each analysis joins two large datasets β€” Web Traffic (what subscribers browsed online) and Foot Traffic (where subscribers went physically). The join reveals the complete consumer journey: "did someone browse Domino's on their phone and then visit a pizza place this week?"

The input schemas are richer than a basic session log. The Web Traffic schema includes cell_id (the serving radio cell β€” provides network-topology geographic resolution without storing GPS coordinates), bytes_up (upstream bytes β€” elevated for transactional sessions such as bet placements and form submissions), and event_count (distinct URL requests within the session β€” discriminates active browsing from background sync at 0). The Foot Traffic schema includes precise dwell window bounds (ping_ts start + end_ts end β€” no estimated dwell_minutes required), cell_site_id (network anchor for geographic fallback when GPS is unavailable), and pre-computed POI matching (poi_id, poi_category: retail Β· transport Β· leisure Β· health Β· hospitality Β· office) resolved at ingestion via H3 spatial join.

The raw data volumes are large: 4 TB/day of web traffic and 10 TB/day of foot traffic. But a key optimisation is that BigQuery only scans the columns needed for a specific question (column pruning). The model assumes 35% of columns are needed on average, which reduces the effective scan volume significantly.

The cost calculator has four sliders: Months of data to include in each analysis (more months = more history but higher cost), Analyses per year, Output dataset size (how much result data is written to BQ after each run), and the GCP discount.

There is no input storage cost β€” the web and foot traffic data already exists in the ingestion pipeline (it's not stored separately for Omni Insights). Only the output β€” the results of the analysis β€” is stored and charged.

Strategy: BQ On-Demand is the right pricing model for ad-hoc work because the frequency is unpredictable. If analyses were run daily, a slot reservation would be cheaper β€” but at 12 per year (monthly cadence), the overhead of maintaining a slot reservation outweighs the savings. The "no input storage" principle avoids double-charging: the raw data is already paid for by the ingestion pipeline.
πŸ“Key assumptions
  • Web Traffic: 4 TB/day raw β†’ 34.3 TB/month compressed (Parquet at 3.5Γ— compression, ~80 bytes/row uncompressed). Foot Traffic: 10 TB/day raw β†’ 85.7 TB/month compressed (~50 bytes/row delta-encoded). Total: 120 TB/month combined.
  • Input schema enrichments: Web Traffic carries bytes_up (upstream volume), cell_id (serving radio cell), and event_count (SNI request count). Foot Traffic carries precise dwell window bounds (ping_ts + end_ts), cell_site_id (network anchor), and pre-matched POI fields (poi_id, poi_category: 6 venue types). These fields expand the analysis surface without changing scan volumes.
  • Column pruning at 35% β€” only 35% of columns are scanned per analysis on average. This is a conservative efficiency assumption.
  • Join overhead: 15% β€” joining two large datasets requires BQ to process slightly more data than a single-table scan.
  • BQ On-Demand rate: $6.25/TB scanned (GCP list price, July 2026).
  • Output: 50 GB/analysis written to BQ. Active storage ($0.02/GB/month) for recent results, long-term ($0.01/GB/month) for older ones.
  • Default: 3 months of data, 12 analyses/year β€” one analysis per month covering the last quarter of history.
Caveats: The column pruning efficiency (35%) is an estimate β€” actual savings depend heavily on how the analysis SQL is written. A poorly written query that selects * would scan all columns and cost ~3Γ— more. The output dataset size (50 GB) is also an assumption β€” if results are written at full row-level granularity (per-subscriber, per-event), this could be many times larger. Cost scales linearly with both the number of months of data included and the number of analyses run per year.

πŸ“¦ Raw Storage Forecast

The foundation β€” storing 20 TB/day of network signals for 5 years

Before any product can run, the raw network data has to be stored somewhere. This page models exactly how much that costs over five years: 20 terabytes of raw telecom signals arrive every day, and they flow through two storage tiers β€” a "hot" BigQuery layer for recent data, and a permanent cold archive in GCS for older data that you rarely need to touch.

5-Yr Storage Total
$822.7K
BQ + GCS Archive
Daily Ingest
20 TB
Fixed rate
2026 H2 Cost
$44.2K
Jul – Dec 2026
By 2030
~$20K
Per month
πŸ“–What this page shows

Data arrives at 20 TB/day and is stored in two places simultaneously. The BigQuery hot tier keeps the most recent 6 months of data in a query-ready format (BigQuery Capacitor columnar format, roughly 5Γ— compression). This is what the daily pipelines scan β€” it's fast but costs $0.02/GB/month.

Data older than 6 months flows to GCS Archive β€” Google's cheapest storage class at $0.0012/GB/month (about 17Γ— cheaper than BigQuery active storage). It's stored as compressed Parquet files. You can't query it directly from BigQuery without first restoring it, but retrieval costs $0.05/GB if you ever need it. Archive storage accumulates permanently β€” nothing is deleted.

The forecast is interactive: you can change the daily ingestion rate (5–100 TB), the BigQuery rolling window (3–24 months), compression ratios, and egress percentage. The 72-month stacked chart updates in real time to show how each change affects the 5-year total.

Strategy: The two-tier architecture is the standard cost-optimisation pattern for large-scale data platforms. Keeping everything in BigQuery would cost ~17Γ— more for data older than 6 months. GCS Archive is practically free per GB but the total volume accumulates β€” 20 TB/day Γ— 1,825 days = ~36 petabytes of raw data by end of 2030 before compression. The model uses 3.5Γ— Parquet compression for archive, bringing that to ~10 PB compressed, at ~$12K/month by 2030.
πŸ“Key assumptions
  • 20 TB/day fixed ingestion β€” no year-on-year growth assumed for raw data volume (network capacity is relatively stable; subscriber count grows but individual usage per subscriber has diminishing marginal volume).
  • BigQuery rolling window: 6 months (configurable 3–24 months). Data older than this window is migrated to GCS Archive.
  • BQ Capacitor compression: 5Γ— β€” the native columnar format compresses well for telecom signal data. 20 TB raw β†’ 4 TB in BQ.
  • GCS Archive compression: 3.5Γ— using Parquet + Snappy codec. 20 TB raw β†’ 5.7 TB in archive per day.
  • 5% monthly egress from BigQuery β€” representing queries that pull data out of GCP (e.g., to on-premise systems or external dashboards). Egress costs $0.085/GB.
  • GCP pricing (July 2026): BQ active $0.020/GB/month, BQ long-term $0.010/GB/month, GCS Archive $0.0012/GB/month.
Caveats: GCS Archive retrieval costs $0.05/GB β€” if historical data needs to be reprocessed frequently, this adds up quickly. The model treats egress as a flat percentage of BQ tier volume; actual egress depends entirely on how downstream systems are built. The 5-year total of $822.7K is dominated by GCS Archive accumulation in years 3–5, not by the per-month BQ hot tier. If the data retention policy changes (e.g. regulatory requirement to delete after 2 years), the archive costs drop dramatically.

πŸ“‹ Data Specification

What data goes in β€” field by field

This page documents the input datasets that feed all Eureka products. It describes the raw network signals at the field level β€” what each column means, its data type, its volume, and how it gets transformed as it moves through the pipeline. It's the ground truth reference for platform engineers and data scientists.

Web Traffic
~58 GB
Per day compressed
Foot Traffic
~37 GB
Per day compressed
Web Events
1.25B
Rows per day
Foot Events
2.5B
After deduplication
πŸ“–What this page shows

The Data Specification covers three source datasets. Web Traffic (network-observed browsing events): each row is a deduplicated session where a subscriber's device connected to a web domain β€” it captures the hostname, session duration, upload/download volumes, and a subscriber token. Foot Traffic (cell site attachment events): each row is a location ping derived from which cell tower a subscriber's device attached to, with dwell time estimation. Cell Site Master: the reference table that maps cell tower IDs to real-world geographic coordinates and coverage areas.

The specification includes aggregation rules β€” how raw events are rolled up from individual session-level rows into 5-minute and 15-minute summaries that the pipelines actually query. This aggregation is what makes the downstream compute costs feasible: without it, every query would scan billions of rows instead of millions.

Strategy: Documenting data at field level serves two purposes: it makes cost assumptions auditable (every cost model references specific fields and their byte sizes), and it ensures that teams building downstream products agree on what the data means before writing a single line of code. The field-level byte counts are used directly in the storage cost calculations β€” e.g. the 58 GB/day web traffic figure comes from summing the byte widths of each field across 1.25B rows and applying Parquet compression.
πŸ“Key assumptions
  • Web Traffic rolling lookback: 6 months β€” BQ partition contains 6 months of daily web events at any point in time.
  • Foot Traffic rolling lookback: 6 months β€” same window as web traffic to support join operations across both datasets.
  • Raw foot traffic: ~8 billion pings/day β†’ 2.5 billion after deduplication. Cell towers generate enormous numbers of attachment events; deduplication collapses multiple pings from the same subscriber at the same site within a time window into a single dwell event.
  • TOKEN_ID links web and foot records β€” a pseudonymous subscriber identifier that connects browsing behaviour to physical location without exposing the actual subscriber ID.
Caveats: Field names shown are operational/internal identifiers β€” they should not appear in any external-facing product UI or documentation. The byte sizes are measured from a sample file; actual production volumes may vary Β±15% depending on network conditions and seasonal usage patterns. The deduplication factor for foot traffic (8B β†’ 2.5B) is an average β€” peak times (commute hours, weekends) may have lower deduplication rates.

πŸ—οΈ Architecture

How all the pieces connect on GCP

The Architecture page shows the end-to-end technical blueprint: how raw network data flows from the operator's source systems, through ingestion and processing layers in GCP, and finally to product outputs and client delivery. It's the map that explains why the cost model looks the way it does.

πŸ“–What this page shows

The architecture is a BigQuery-native data lakehouse: raw on-prem data is filtered, transferred via STS over Interconnect into a restricted GCS bucket, masked by Cloud DLP, and loaded into BigQuery via a BQ Load job. All processing runs inside BigQuery β€” Omni Alpha uses a dedicated 700 BQ Enterprise Slots reservation (5 hr/day, $154/day, 80 TB/day scan across 11 tables); Omni Enterprise uses a separate 500-slot pool (10 hr/day, $220/day, per-category mode); Omni Pinpoint uses BQ Flex Slots (500 slots Γ— 6 hrs/day, $120/day) for the daily H3 spatial join + 4M-row demographic enrichment of 35B foot traffic rows into 1.2M-POI hourly and daily aggregation tables; Omni Insights uses on-demand BQ queries for ad-hoc cross-channel analysis. No Spark, no Dataflow, no separate compute cluster to maintain.

Outputs are delivered through BigQuery Analytics Hub for structured dataset sharing (no data egress, no copying), and via S3 bucket export for scheduled report delivery. Cloud Composer orchestrates every production job β€” DAG sensors, DQ gates, SLA alerts, and slot reservation lifecycle.

Caveats: The architecture diagram represents the planned production state. Some components may be phased in over time rather than deployed simultaneously at launch. Cloud Functions orchestration cost is not currently included in the cost models β€” it is expected to be negligible (<$50/month) at this pipeline frequency. The diagram does not show disaster recovery or multi-region replication, which would add to the storage cost if required.