Kairos — Warehouse Groundwork after the Red Team (2026-08-17)
What the warehouse said on the last live-token evening, after an opposing review broke the loose readings. Two evenings of BigQuery reads (16 Aug, ~425 GB, every result saved in full) produced eight headlines; a findings-only red team on the strategy and the headlines returned ten findings; every test it named that the warehouse could answer was run before anything reached the board (G7, 60 GB). This note carries what survived, in the words that are now allowed. Canonical copies: Kairos repo
CONTEXT/2026-08-16-warehouse-g1…g6*.md(with dated correction blocks),CONTEXT/2026-08-17-warehouse-g7-talos-tests.md,CONTEXT/warehouse/contracts/2026-08-17-critical-table-contracts.md,CONTEXT/warehouse/MANIFEST.md. If this note and the repo disagree, the repo wins.
What is now measured (L2, working tier)
- Churn by country tier, on a month-start cohort. Users entitled at the prior month-end, followed to their first churn, paired to reactivation within 30 days. Feb–Jul 2026: tier 1 gross 24.6–34.6% / net-30 18.8–25.7% · tier 2 23.5–31.2% / 18.0–24.1% · tier 3 38.2–55.5% / 28.4–43.1%. July: 24.6 / 18.8 · 23.5 / 18.0 · 38.2 / 28.4. Tier 3 turns over about 1.5× tier 1, every month. The "≈3-month average life" reading is a hypothesis until a survival curve by plan duration is drawn (yearly and 2-year plans are a tier-1 phenomenon).
- One list price across tiers on the direct gateways. July clean new purchases,
plan_monthly_plus: modal $13.49 in every tier on Stripe (447 of 448 tier-1 purchases, 240 of 257 tier-3), Coingate, PayPal, Apple. Google Play tier 3 lands at $12.55 modal, dispersed — store-local pricing, not an MN policy. Nigeria is the #1 country by 2026 new-purchase count (8,208), on Google Play at $12.90 a month. - Tier-3 cancellers name price less and product failure more — on every platform. 2026 in-app cancel submissions, standardised within platform: too-expensive Android 22.6% (tier 3) vs 29.9% (tier 1); iOS 19.3% vs 28.6%; Windows 15.6% vs 24.9%; disconnects, blocked sites, missing features, usability and error 7040 higher in tier 3 on each. This is a self-selected survey sample (~7.5k submissions against ~44k churn rows, app-only); it is a direction, not a churn cause, and it cannot rule price out.
- About 48% of 2026 churned subscriptions carry a billing-retry label; 52% voluntary, including explicit cancellation requests. In the labelled table one churn row is one subscription and one user (31,398 rows / 31,394 subscriptions). About a quarter of churn rows in the status ledger reverse within 30 days, and reactivations inside 30 days are almost all recurring-type — billing recoveries, not returning customers. Whether a labelled churn was involuntary in outcome, and how large the never-recovered remainder is, needs MN's pairing key (ask 14).
- Primer replaced Stripe on the web in the week of 22 July 2026 and the revenue spine still labels it Stripe. Reconciled weekly on identical New-type populations, web dollars ran at 101 · 102 · 101 · 100 · 100 · 110% of the spine's across the switch — no aggregate discontinuity. Row-level completeness cannot be checked from the warehouse (the web table has no transaction id). Any Stripe-keyed analysis silently absorbs Primer traffic.
- The iOS first-connect gap was composition, not product. Among users entitled to connect at their first try, same-day success is 96.9% on Android and 96.3% on iOS. Unentitled or null-status first-tries are 52% of the iOS cohort and 20% of Android; they connect ~8% of the time on both. The candidate row was dropped before release.
- An api_error does not lower a session's chance of connecting. Sessions with an api_error that tried to connect logged
connect_success95.6–96.0% (Android) / 91.5–93.7% (iOS) of the time, against 94.3–95.7% for clean sessions; 81–88% of sessions with an error on the connect endpoint itself connect after the error. Tunnel time per session cannot be tested (the duration table has no session key). A backend/product note, not a Kairos row. - The warehouse, measured (2026-08-17 00:44, first clean inventory).
mysterium-bqholds 24 datasets, 503 tables, ~980 GB logical:App_usage642 GB (38 tables, 2.23 bn rows) · GA4 exportanalytics_354300879153 GB (180 daily shards) ·airbyte_internal150 GB ·postgre21.6 GB (66 tables) ·airbyte3.3 ·looker2.7 ·claude1.8 · everything else under 2 GB each (Marketingis one table of 38,604 rows). Both automated inventories on 16 Aug had reported rows=0 / GB=0 for every dataset — the storage query was the wrong form for this region and every dataset fell to a name-only fallback; corrected and re-run while the token lived (KairosCONTEXT/warehouse/2026-08-16T214427Z-inventory-*, 0 errors). Two readings follow: the two evenings' ~485 GB of reads are not "3 of 24 datasets" by weight — the churn/payment spine (postgre,claude,looker, ~26 GB) has been read many times over,App_usage,airbyteandimportswere opened for G1/G3/G7 (four of the 38App_usagetables are under contract), and the untouched mass is the rest ofApp_usageand the GA4 export; and the "3% of what happens in MN" hunch cannot be scored from bytes alone, because the heaviest datasets are event logs, not distinct kinds of fact. - The revenue spine moves a little every night; the clean table does not (2026-08-17 01:15). The 16 Aug batch happened to read the same tables twice, seven hours apart, across a nightly rebuild.
claude.vpn_purchasescame back identical on every one of 151 historical month × type cells, rows and cents; the Stripe charge replica likewise, except for the live month growing.postgre.mv_user_revenue_actualand its BI copylooker.users_revenuedid not: 103 of 227 historical cells (2023-04 → 2026-07) moved between the two reads — never by more than 7 rows or $174 in a cell, net −14 rows and −$201 across the whole history — mostly inside the duplicate zero-length rows, but 25 cells inside real dollars too, typically a relabel between New and Recurring in the same month (Jan 2026: New +$13.24, Recurring −$13.24). July 2026 real revenue on the spine read $162,517.60 at 14:03 and $162,531.21 at 21:27; the clean table read $162,517.60 both times. Rule from here: a spine figure carries its rebuild stamp, two spine reads that differ by tens of dollars are rebuild noise not late data, and the clean table is the one that can be quoted twice. Ask 3 to the revenue-model owner now includes what orders the type derivation. KairosCONTEXT/2026-08-16-warehouse-governance-findings.md§ G, items 30–31 (31: the Stripe attempt-level success share has been flat at 23–30% by count for fifteen months, through the March 2026 Adyen→Stripe step — a dunning change would show in that series as a step that has not happened). - The "3% of what happens in MN" hunch, measured two ways (01:43). Bytes cannot score it. First cut: every workstream MN's leadership named across the 68 all-hands transcripts (2021–2024) plus the 08-07 deck — 306 entries, merged into 32 kinds of fact, weighted by mentions × airtime — against the measured 503-table inventory: about a third (30% weighted, 34% on 2023–2026, 44% unweighted). Its shape: the subscription-database part of the VPN (revenue, churn, base, usage, payments, support) reads 84–89% present; proxy + B2B 13%; nodes + token + protocol 17–19%; company functions 12–16%. Lee's correction, same hour: that VPN row was drawn around the database and then scored as full — marketing, ads, funnels, SEO, content, store pages, experiments and user research are the VPN business too, and none of it is in there. Second cut, VPN only, by capturability: of 36 information streams a consumer-VPN business runs on, the warehouse holds 11 (~30%); 14 more exist at MN outside it and sit ungranted on the data-access board (ad accounts and spend, GA4 history, SEO tool, affiliate payouts, store consoles and reviews, product analytics, error tracking, A/B tests, PayPal/Coingate charges, fraud tooling, pricing tables, Slack/ClickUp/Notion, P&L); 7 exist nowhere at MN today (metric dictionary, experiment register, creative-level ad data, channel-cohort LTV, churn interviews, per-feature usage, brand demand); 4 are not capturable (the non-buyers, market and platform shifts before they hit the numbers, sentiment beyond exports, tacit knowledge). Above all of these sits the soft-data layer Lee named — the power-customer avatar, the acquisition journey, triggers, objections, positioning: in 68 all-hands and the 08-07 deck the VPN has no persona, ICP or journey mention at all (the proxy side has one), and the deck's own top priority, "high-quality customers", is undefined; the warehouse can sketch the payer and the last third of the journey, not the person or the first two thirds. Present is not trustworthy, and a paying-base database says nothing about how people came to pay or why others did not. On the raw-events denominator nobody can measure, 3% may still be right. Kairos
CONTEXT/2026-08-17-warehouse-coverage-vs-what-mn-does.md(both cuts, limits, per-entry evidence inCONTEXT/warehouse/coverage/).
What was withdrawn or narrowed, and why
- "iOS converts 4× Android" (withdrawn 16 Aug — instrumentation step at app 2.4.4) · "West Africa iOS pays" (withdrawn 16 Aug — the money says Nigeria on Google Play, Côte d'Ivoire on Apple) · "99% of error sessions connect" (arithmetic artefact → 96/92) · "no revenue missing" through the Primer switch (mixed populations → the identical-population reconciliation above and its limit) · "the base grows on retention, not acquisition" (no 2025 H1 in the ledger; the causal sentence is gone, the fact stands: base +15% Jan→Jul while new subscriptions fell 27% and monthly churned users fell from 6.4–8.2k in autumn 2025 to 5.1–5.9k in 2026) · "tier-3 churn is product, not price" (a survey direction, kept only in that form).
What it changed on the board
/opportunities/ v0.6.0 (live 17 Aug 00:20): three warehouse-lane rows added and signalled — W6 Android product experience in the largest paying geography · W7 failed payments as a churn label and the never-recovered remainder · W8 the Primer switch and the payment label it hides under; W3 pricing reworded (tier churn is not elasticity evidence; the lever is product); Retention Economics (#02) gains its measured tier cut as a sixth graded claim. The iOS row never shipped.
Asks that only MN can answer
14 — the key that pairs a labelled churn with its billing attempts and 30-day recovery · 15 (owner half) — churn rate by tier on MN's own definition · 16 — the Primer transaction identifier and where it lands in the spine · 17 — any 2025 H1 subscription series for the same-month comparison · Play Console grace-period setting history (the IN_GRACE_PERIOD → 0 signal stays a hypothesis) · the durable BigQuery principal — every job from G7 onward carries Kairos labels; the ~300 before it do not, and that is said plainly in the ask.
The method lesson (Talos F1, confirmed)
Broad capture while a token lives was right; drawing use-specific conclusions from tables whose grain, identity, date field and joins had not been contracted was not — it produced the withdrawn headlines. The six critical tables now have written contracts; a claim outside a contract is working-tier until the contract is extended.