Skip to content

Audience Persona & LTV Indexing Methodology

The Base Audience Intelligence Engine transforms raw customer transaction data into empirical psychographic segmentation, demographic baselines, and commercial value distributions.

This document explains the mathematical foundations, statistical guardrails, and strategic formulas used to evaluate and rank audience personas.


1. The 1:1 Identity Principle

A foundational principle of the Base intelligence engine is that all persona assignments, indexing scores, demographic profiles, and LTV valuations are derived strictly and exclusively from 1:1 exact email matches (real, identified individuals).

graph TD
    A[Raw Customer Upload] --> B{Exact 1:1 Email Match?<br/>responseType = 'xEmail'}
    B -- Yes (Real Person Match) --> C[80 Psychographic Personas]
    B -- No (Unmatched / Missing) --> D[Excluded from Persona Denominator<br/>Preserved in Audit Log]

    C --> E[Audience Share %]
    C --> F[Audience Index Score %]
    C --> G[LTV Value Ratio vs Mean]

    E & F & G --> H[Materialized Audience Mart]
    H --> I[Executive Strategic Brief & Creative Directives]

Why Exact Email Matches Are Mandatory

Traditional market research often relies on geographic proxies (such as ZIP code centroids or census tract averages). However, a single ZIP code typically contains dozens of conflicting lifestylesβ€”from young single renters to affluent retirees.

Averaging ZIP codes introduces geographic blurring that destroys psychographic signals. Base guarantees 100% person-level integrity by requiring cryptographic SHA-256 email resolution for all persona assignments.


2. Core Mathematical Formulations

Every snapshot upload materializes standard statistical metrics across all 80 personas based solely on exact email matches:

A. Audience Share (%)

The percentage of a brand's verified email-matched customers represented by a given persona:

$$\text{Audience Share} = \left( \frac{\text{Email Matched Count for Segment}}{\text{Total Email Matched Audience Count}} \right) \times 100$$

  • Interpretation: If a brand has 10,000 email-matched customers and 1,200 belong to #MidasMight (Segment A01), the Audience Share for A01 is 12.0%.
  • Denominator Integrity: Only records with confirmed responseType = 'xEmail' and valid segment codes ($\ne \text{'Z00'}$) enter the calculation.

B. Audience Index Score (National Lift)

The Audience Index Score measures how heavily a brand's verified customer base over-indexes or under-indexes relative to the natural United States population baseline:

$$\text{Audience Index Score} = \left( \frac{\text{Audience Share \%}}{\text{National Population Share \%}} \times 100 \right) - 100$$

Index Classification Thresholds:

Index Score Range Status Strategic Meaning
$\ge +20.0\%$ OVER_INDEXING Core Audience Segment: Significant concentration compared to national baseline. High brand affinity.
$-20.0\% < \text{Score} < +20.0\%$ NEUTRAL Baseline Representation: Appears in normal proportion to the general population.
$\le -20.0\%$ UNDER_INDEXING Under-Represented / Averse: Rarely engages with or purchases the brand's offering.

Example Calculation

  • Persona: #SuburbanSuccess (Segment B02)
  • Brand Audience Share (Email-Matched): $6.0\%$
  • National Population Share: $2.5\%$ $$\text{Index Score} = \left( \frac{6.0}{2.5} \times 100 \right) - 100 = 240 - 100 = \mathbf{+140.0\%}$$
  • Result: OVER_INDEXING ($+140.0\%$ lift over national baseline). This segment is 2.4Γ— more prevalent in your customer base than in the US general population.

C. Persona-Level LTV Ratio vs Audience Mean (ltv_vs_mean)

To determine whether a specific persona spends more or less than the brand's average customer, Base calculates the relative monetary ratio across verified individual buyers:

$$\text{LTV Index} = \frac{\text{Segment Mean Customer Value}}{\text{Overall Audience Mean Customer Value}}$$

  • 1.00: Spends exactly equal to the brand average.
  • > 1.00 (e.g. 1.45): Spends 45% more than the average customer.
  • < 1.00 (e.g. 0.78): Spends 22% less than the average customer.

πŸ›‘οΈ Statistical Guardrail ($N \ge 30$)

To prevent low-volume sample sizes from introducing statistical noise or distorting commercial strategy, Base applies a strict minimum threshold:

$$\text{ltv_vs_mean} = \begin{cases} \text{ROUND}\left(\frac{\text{Segment Mean LTV}}{\text{Audience Mean LTV}}, 2\right) & \text{if } \text{Segment Email Matched Count} \ge 30 \ \text{NULL} & \text{if } \text{Segment Email Matched Count} < 30 \end{cases}$$

Why Guardrails Matter

If only 2 customers in a list of 50,000 belong to a rare segment and one happens to place a large $5,000 order, reporting an LTV ratio of 8.5Γ— would create a misleading incentive to invest heavily in an unproven segment. Suppressing ratios for $N < 30$ ensures decisions are backed by statistical significance.


3. The 4-Quadrant Strategic Actionability Matrix

By evaluating Audience Index Score (affinity) alongside LTV vs Mean (monetary value), marketers and autonomous AI agents can instantly categorize personas into actionable growth strategies:

                      High LTV (> 1.10x)
                             β–²
                             β”‚
     [Quadrant 3]            β”‚            [Quadrant 1]
   HIDDEN HIGH-VALUE         β”‚            CORE WHALES
  Niche premium buyers;      β”‚       High affinity + top spend;
  tailored messaging only    β”‚       primary lookalike seed
                             β”‚
 ◄───────────────────────────┼───────────────────────────► High Lift
 Low Lift (< -20%)           β”‚                             (> +20%)
                             β”‚
     [Quadrant 4]            β”‚            [Quadrant 2]
    NEGATIVE TARGETS         β”‚         VOLUME FOUNDATION
  Under-indexing & low value;β”‚       High brand affinity,
  exclude from paid ads      β”‚       baseline transaction value
                             β”‚
                             β–Ό
                      Low LTV (< 0.90x)

Strategic Applications:

  1. Quadrant 1 (Core Whales β€” High Index, High LTV):
  2. Action: Seed audience for programmatic lookalikes, meta-audiences, and high-touch VIP retention campaigns.
  3. Quadrant 2 (Volume Foundation β€” High Index, Average/Low LTV):
  4. Action: Core messaging for broad brand awareness and mainstream customer acquisition efficiency.
  5. Quadrant 3 (Hidden High-Value β€” Low Index, High LTV):
  6. Action: High-margin micro-campaigns with specialized messaging; do not dilute broad marketing to target them, but nurture them when they arrive.
  7. Quadrant 4 (Negative Targets β€” Low Index, Low LTV):
  8. Action: Add as explicit negative exclusions in Google Ads / Meta Ads to eliminate ad spend waste.

4. Demographic Cross-Rollup Methodology

In addition to psychographic classifications, Base automatically projects 70 empirical demographic attributes (such as Median Age, Household Income, Home Ownership %, Education %, and Commute Times) without requiring invasive 1:1 PII harvesting.

How it Works:

Each of the 80 Spatial Personas carries a verified national demographic vector derived from US Census Bureau and American Community Survey (ACS) datasets.

When a customer snapshot is analyzed, Base computes the dot product of the 1:1 Email-Matched Audience Share distribution against the Demographic Matrix:

$$\text{Metric}{\text{Audience}} = \sum_i \right)$$}^{80} \left( \text{Audience Share}_i \times \text{Metric Baseline

This delivers: * Exact projected household income and age distributions. * College degree percentages and home equity insights. * Longitudinal shifts in customer wealth across quarterly snapshot comparisons.


5. Multi-Source Identity Resolution & Cache Hierarchy

Customer records are matched to exact personas through a strict cryptographic pipeline:

graph LR
    Input[Incoming Customer Email Hash] --> CheckCache{Match in<br/>email_hash_segments?}
    CheckCache -- Yes (Cache Hit) --> Hit[Instant 0-Cost Match<br/>responseType = 'xEmail']
    CheckCache -- No (Cache Miss) --> SpatialAPI[Spatial Persona API<br/>Deterministic Match]
    SpatialAPI -- Resolved --> Enriched[responseType = 'xEmail'<br/>Persisted to Cache Table]
    SpatialAPI -- Unresolved --> Unmatched[responseType = 'UNMATCHED'<br/>Excluded from Personas]
  1. Instant Identity Cache (email_hash_segments):
  2. Pre-resolved, zero-cost matching against our growing global identity cache table.
  3. Live Deterministic API Resolution:
  4. For cache misses, individual SHA-256 email hashes are queried against the Spatial persona identity engine.
  5. Compounding Cache Harvester:
  6. All newly resolved email matches across customer uploads and real-time website leads are continuously synchronized into the identity cache, accelerating all future runs across the platform.