☿ Kairos War Room
Gated source note · private

Autonomous agent relay — strongest use-case portfolio

[!summary] Verdict The relay is most valuable where two or more independent principals hold different private context, the handoff itself is expensive, and the result can be verified. It is not most valuable for generic agent teamwork, polling, messaging, or raw model capacity.

Inside MN/Kairos, the strongest use is a cross-boundary opportunity and exception engine: local agents interrogate the sources they are authorized to see, exchange bounded evidence rather than raw private data, independently attack conclusions, and return only accepted artifacts or genuine human decisions. The nearest cash wedges are payment recovery and GoProxies paid-test conversion; the largest compounding wedge is the Standard of Performance ratchet across vital operations.

Outside MN, the strongest commercial form is a managed vertical outcome service, not a relay platform: cross-company billing and commercial exceptions, third-party cyber incident rooms, healthcare revenue-cycle exceptions, and M&A/JV diligence. Sell resolution, recovery, assurance, or time saved. Do not sell “agents talking to agents.”

Generic identity, governance, payments, legal context, agent coordination, and certification are filling rapidly with standards and large-platform products. The defensible layer is the vertical workflow, local context, acceptance test, evidence receipt, and named failure owner.

Evidence state

This is a deep strategic synthesis current to 2026-08-20, not an independently reviewed recommendation. It uses the existing vault research rather than repeating it, then adds a narrow blind-spot scan of cross-organizational cost pools, emerging vertical competitors, and current agent reliability.

Current runtime reality matters. At 2026-08-20 16:52 EEST, agent-relay reported 40/40 daily hops consumed, 77 unread Codex messages, 12 unread Claude messages and 16 parked items. The code supports read-only judgment hops, typed dispositions, durable replies, review-request priority, non-quota retry parking and canon-change review initiation. Its isolated red suite passed 85/85, but case 8 proves the live quota defect rather than closing it: a quota failure deliberately stays unparked and retries on the next run, bypassing MAX_ATTEMPTS. That matches the measured 66 attempts on one review request. The backlog, exhausted cap and unbounded cross-run quota retry mean the relay is not production-load-ready and should not be sold or relied upon as a high-availability service yet.

The opposing reviewer is Claude/Tris. Tris is unavailable on provider quota until 2026-08-22 15:00 EEST, so no independent verdict is incorporated here. Metis cannot substitute for private Lee-vault judgment: he serves another principal and can inspect only a deliberately shared-safe packet.

Source-class coverage

Source class Status Use here
Current relay contract, code and live status found establishes actual capability and present bottlenecks
Existing Kairos agent/context-transfer research found avoids repeating A2A, MCP, identity, clean-room and competence-transfer work
Existing agent-market, credit-market and provider-terms research found rules out public allowance resale, generic marketplaces and protocol-first entry
Current MN economics and opportunity corpus found anchors the internal ranking in payment recovery, paid tests, warehouse coverage and the September mandate
New cross-organizational cost-pool evidence found commercial disputes, healthcare administration, third-party cyber and M&A diligence
New direct competitors and substituting infrastructure found billing-dispute multi-agent work, enterprise agent identity/governance and legal-context standards
Independent buyer demand for this exact relay composition not found remains a market hypothesis; adjacent demand does not prove product demand
Provider permission or legal opinion for consumer-plan resale not obtained public allowance exchange remains blocked
Opposing-agent strategic review blocked by current reviewer quota analysis remains writer judgment, not reciprocal closure

The relay-fit test

Use the relay only when all six conditions hold:

  1. There is a real boundary. Two principals, companies, private corpora, authority domains, or vendors possess information that should not be pooled casually.
  2. The boundary is economically expensive. Human relaying, repeated re-briefing, missing context, slow dispute handling, or mistrust materially delays revenue, increases loss, or consumes senior attention.
  3. The task ends in a checkable object. A reconciled ledger, evidence packet, test result, sourced decision, accepted artifact, incident timeline, or bounded recommendation exists.
  4. Authority can be made explicit. Read, propose, execute, spend, disclose, settle, deploy and communicate are separate permissions.
  5. Failure has an owner and a stop. The workflow names who absorbs a bad result, when the agent must escalate, and how authority is revoked.
  6. The value at risk dwarfs review cost. A useful default is expected value at least ten times the coordination and verification burden.

If a single agent can access the complete authorized corpus and produce the same verifiable result, use one agent. The vault already carries independent evidence that multi-agent boundaries can lose information; multiplying agents without a boundary-specific reason is negative ROI.

A compact economic test:

relay value = value at risk × probability the relay changes the outcome × capture rate − execution − verification − dispute − regulatory cost

The relay earns its place by changing the outcome or reducing human attention. Message count, agent count and token use are not value metrics.

A. Within MN/Kairos

Scores below are comparative judgment on a 1–5 scale, not measured economics. Total /20 is the equal-weight sum of upside, relay-specific fit, current evidence and time to proof. Gap shows distance from the section leader. Strategic rank remains separate because readiness and authority constraints are qualitative rather than smuggled into the arithmetic.

Strategic rank Use case Upside Relay-specific fit Evidence now Time to proof Total /20 Gap Readiness
1 Cross-boundary opportunity engine 5 5 3 4 17 0 pilotable after control-plane repair
2 Revenue and payment exception room 4 4 4 5 17 0 nearest measurable cash test
3 GoProxies paid-test conversion bench 4 5 3 4 16 −1 gated on account-level access and owners
4 Standard of Performance ratchet 5 4 3 3 15 −2 method defined; source access incomplete
5 Cross-company incident and dependency room 4 5 2 4 15 −2 tabletop-ready, production authority absent
6 Context continuity and competence ladder 3 5 4 3 15 −2 integrity shown; retrieval gate still open
7 Allowance sweeper and burst research queue 2 direct / 4 option 1 2 5 10 direct / 12 option −7 / −5 useful experiment, weak relay moat

1. Cross-boundary opportunity engine

This is the strongest total use because it turns the relay into an information-arbitrage machine without collapsing the principals' private worlds into one store.

The agent beside each source answers only what that principal can authorize. A Kairos agent can hold the market map, opportunity portfolio, external benchmarks and kill tests. An MN-situated agent can hold customer, support, payments, product, infrastructure or commercial records. They exchange bounded questions, aggregates, contradictions, evidence receipts and next falsifiers. A third pass attacks whether the conclusion follows. Humans receive only a real authority decision, not the integration work.

This is the relay form of the existing Kairos move: measure something the warehouse cannot see, pair it with something it can, and test the join. It applies to:

The financial upside is option-like rather than presently quantifiable. It can kill a false opportunity before material build or surface an opportunity that no single data source can reveal. Its asymmetry comes from low-cost repeated falsification against potentially large strategic decisions.

First falsifier: select three current opportunity questions whose decisive evidence is split across Lee-side and MN-side context. Require each side to answer from local sources without moving raw primary data. Measure time to a source-cited decision, human-attention minutes, first-pass acceptance and whether the relay changed the ranking. Kill the claim if the same result is cheaper from one agent with a bounded export.

2. Revenue and payment exception room

This is the nearest measurable cash wedge.

The current payment corpus identifies an approximately $18.5k monthly July Stripe involuntary renewal pool after matched cancellation analysis. That is unpaid face value, not recoverable revenue. Scenario arithmetic shows the scale without pretending to forecast it:

Incremental recovery of the measured monthly pool Monthly gross recovered Annualized gross recovered
5% $925 $11.1k
10% $1.85k $22.2k
20% $3.7k $44.4k

The current process already retries heavily, so the relay does not earn value by saying “retry more.” It would join four private contexts that otherwise fragment the decision: invoice-grain evidence, decline-class/payment-processor evidence, customer intent/support evidence, and the product owner's reversible intervention path. The output is a pre-registered eligible cohort, proposed routing/timing change, guardrails and accepted measurement—not a dashboard.

The same room can later handle revenue-perimeter contradictions, duplicate-carrying tables, processor reconciliation and support-linked churn exceptions.

First falsifier: replay 20 historical failed-renewal exceptions across two bounded agent contexts. Require a reproducible eligibility decision and evidence citation for each. Compare against a single-agent baseline. Proceed only if the relay reduces human adjudication time and does not increase classification error or private-data movement.

3. GoProxies paid-test conversion bench

The July company-reported proxy result contained $11,628, of which $9,084—78%—was paid tests. This is paid intent, not recurring revenue. The longitudinal corpus shows the recurring failure: the product became commercially real, while test-to-contract conversion and durable MRR acquisition remained unresolved.

The relay-specific opportunity is not another CRM. It is a persistent case room per test/account in which:

The economic formula is simple: annualized contribution = converted MRR × 12 × contribution margin. Current evidence does not license a conversion-rate or margin assumption.

First falsifier: reconcile every July paid test into one bounded schema and measure the fraction with a reachable economic buyer, repeat need, technically successful test, positive-contribution path, named owner and decision date. Kill the relay layer if it merely restates fields that one account owner already holds.

4. Standard of Performance ratchet

The existing method is: map one vital operation, assemble the specific gold standard, close one 80/20 gap as a non-reverting habit or automation, then repeat.

The relay makes this stronger when each operation crosses functions. One situated agent maps the actual operation from local systems. Another brings the external gold standard and attacks whether it transfers. A verifier confirms that the ratchet changed behavior and did not revert. That is a better use of multiple agents than parallel brainstorming because each lane has a distinct source and acceptance test.

Likely first operations:

  1. analytical-fitness guard before a metric reaches a decision;
  2. paid-test to contract;
  3. failed-payment classification and intervention;
  4. support issue to product change to measured recurrence;
  5. node/supply quality to sellable proxy SKU economics.

The upside is compounding EBITDA and decision quality rather than one visible jackpot. The risk is bureaucracy: if the relay produces more artifacts than changed operations, shrink it.

5. Cross-company incident and dependency room

Payments, cloud/data providers, app stores, infrastructure vendors, proxy suppliers and partners can each create incidents where no party may expose its whole environment. The relay can maintain a fast liveness rail and a durable evidence rail while each party's agent answers bounded questions locally.

The useful output is a shared incident spine: confirmed facts, affected perimeter, hypotheses, tests, owners, decisions, evidence receipts and a closure criterion. Private logs, credentials, customer rows and personnel reads remain local.

The financial payoff is avoided downtime, revenue loss, repeated incident labor and relationship damage. It is event-driven and therefore hard to size from the current corpus.

First falsifier: run a two-principal tabletop on a payment-provider or data-pipeline incident using synthetic evidence. Red-team prompt injection, stale state, conflicting definitions, partial acknowledgement, duplicate ownership and one party going offline. The relay fails if the humans still have to reconstruct the timeline or manually relay every test.

6. Context continuity and competence ladder

This is the most mature conceptual use and an enabling layer for the others.

The current Metis transfer has self-reported 673/673 hash equality with zero mismatches. That proves reported storage integrity, not working competence. The six blind retrieval probes remain the live gate. The stronger product is a ladder:

identity → bounded corpus → integrity receipt → hidden retrieval/boundary probes → reversible work → accepted results → wider authority

This reduces key-person loss, provider-limit interruption, model replacement cost and onboarding time. It also prevents a dangerous equivalence: “the agent has the files” does not mean “the agent understands the local traps” or “the agent may act.”

The direct financial upside is protective and should be measured as reduced human re-briefing, faster recovery and avoided error—not as a standalone certification fee inside Kairos.

7. Allowance sweeper and burst research queue

The existing market research correctly sequences local sweep before exchange: allocate expiring allowance to the principal's own pre-cleared backlog, measure accepted value, and only then test a trusted pool.

Within Kairos the backlog can include competitor diffs, public-surface monitoring, source extraction, historical reconciliation, deterministic QA, red-team packets and bounded alpha-hunt branches. This can create cheap option value and reduce subscription waste.

It is not a strong relay use case by itself. Quota sensing, backlog scoring, scheduling and deterministic checks belong in code. The relay enters only when a task must cross to another principal or receive an opposing judgment. Treat the sweeper as a supply experiment and control-plane feature, not the product or moat.

B. Outside MN/Kairos

There are two rankings: largest economic option and best initial paid wedge. They are not the same. Total /20 is the equal-weight sum of the four numeric dimensions; competitive pressure remains visible but is not assigned a false-precision number. Gap shows distance from the 18-point leader.

Strategic rank Use case Economic option Relay-specific fit Evidence Time to paid proof Total /20 Gap Competitive pressure
1 Cross-company billing and commercial exception resolution 5 5 4 4 18 0 high and rising
2 Third-party cyber incident and vendor-verification rooms 5 5 4 3 17 −1 medium
3 M&A/JV diligence and post-deal integration 5 5 3 3 16 −2 medium/high
4 Healthcare revenue-cycle exception assurance 5 5 5 2 17 −1 high, regulated
5 Verified outcome relay for professional services and software change 4 4 3 5 16 −2 high
6 Organizational succession and managed-service transfer 4 5 3 3 15 −3 medium
7 Private R&D and technical clean-room collaboration 5 5 2 1 13 −5 medium
8 Personal agent portability 4 at scale 4 2 2 12 −6 very high/platform-owned

1. Cross-company billing and commercial exception resolution

The International Chamber of Commerce's 2025 Oxera study estimates that SMEs globally write off about US$1 trillion each year in bad debts and disputed invoices. The missing market exists because evidence collection, negotiation and enforcement can cost more than the claim.

This is almost a perfect relay shape: buyer and seller each hold private ledgers, delivery records, communications and authority. Their agents can reconcile facts locally, identify the exact disagreement, exchange signed evidence, propose a bounded resolution and route only unresolved legal or commercial judgment to people.

But the category is not empty. A 2026 TM Forum project backed by operators and vendors is already demonstrating multi-agent billing-dispute management across invoice, billing, mediation and partner-charge records. The opening is therefore not “AI for disputes.” It must be a narrow buyer wedge, such as cross-border SME invoice exceptions, marketplace seller disputes, channel/partner settlements, telecom wholesale reconciliation or a specific ERP pair.

Business model: managed service plus per-accepted-case or a share of recovered value. Avoid fully autonomous settlement until authority, legal terms and recourse are explicit.

First falsifier: 100 historical closed cases in one vertical. Compare time-to-reconciled-fact-pattern, agreement rate, human minutes and recovered value against the existing process. If most cost sits in unstructured legal negotiation rather than evidence reconciliation, narrow the product to case preparation.

Sources: ICC/Oxera low-value commercial disputes · TM Forum billing-dispute catalyst

2. Third-party cyber incident and vendor-verification rooms

ISC2's 2025 supply-chain survey reports that 28% of respondents' organizations experienced a cyber incident originating from a third-party supplier in the prior two years; the enterprise figure was 34%. The named problems—lack of visibility, inability to verify vendor claims and incomplete knowledge of downstream suppliers—match the relay's cross-principal boundary directly.

A preparedness product can keep each organization's sensitive telemetry local while exchanging current controls, evidence receipts, dependency graphs, incident facts and revocation actions. An incident product can create one auditable timeline without giving either party blanket access to the other's environment.

Business model: recurring preparedness retainer plus incident-response fee. The likely buyer is a security team, MSP, insurer or vendor-risk platform—not an “agent marketplace” customer.

First falsifier: two-company tabletop with a real supplier/customer pair and synthetic secrets. Measure time to shared scope, false claims caught, secrets withheld correctly and decisions requiring human escalation. A static security questionnaire is the baseline to beat.

Source: ISC2 2025 Supply Chain Risk Survey

3. M&A, joint-venture diligence and post-deal integration

This is a high-ACV, low-frequency use. Both sides need to compare facts, definitions, liabilities and operating dependencies without prematurely pooling data. Traditional virtual data rooms move documents; they do not prove that the recipient can interpret local definitions or carry decisions into integration.

The relay can add:

Thomson Reuters' current analysis argues that granular lineage, decision traceability and reusable data workflows are becoming central to diligence and integration. That validates the problem class, not demand for this exact product.

Business model: premium diligence/integration service sold through legal, transaction-services or PE operating partners. Do not start by building a horizontal data room.

Source: Thomson Reuters Institute on data-driven M&A diligence

4. Healthcare revenue-cycle exception assurance

The 2025 CAQH Index, based on more than 600 provider organizations and plans representing 63% of insured lives, estimates a remaining $21 billion annual savings opportunity from further automation of manual and partially manual administrative transactions.

The reliability constraint is equally important. HealthAdminBench evaluates 135 expert-defined prior-authorization, denial/appeal and equipment-order tasks across simulated EHR, payer and fax systems. Its best tested agent achieved only 36.3% full-task success, despite much higher subtask performance. Failures concentrate in long end-to-end coordination.

That argues against autonomous end-to-end replacement today. It argues for a relay that decomposes the workflow, preserves evidence, assigns each system-specific subtask to a situated agent, verifies every checkpoint and escalates the unresolved clinical or coverage judgment.

Business model: exception-resolution and audit service for one narrow transaction class. This has enormous upside and high regulatory/liability burden; it is not the first external pilot.

Sources: 2025 CAQH Index release · HealthAdminBench

5. Verified outcome relay for professional services and software change

This is the best initial paid wedge because it is closest to what the current system has actually done: bounded research, code review, red-team findings, repair requests, evidence receipts and explicit acceptance gates.

Possible buyers include boutique consultancies, dev agencies, security reviewers, compliance teams and managed-service providers. The product is a verified change packet or decision packet with provenance, opposing review, repair state and remaining dissent.

The upside is lower than healthcare or global disputes, but time to proof is shorter and liability can be bounded by choosing reversible, testable work. The weakness is competition: every major agent platform is building orchestration and governance. The wedge must be a vertical acceptance contract and outcome warranty, not a prettier agent inbox.

6. Organizational succession and managed-service transfer

When an employee, agency, MSP or operator leaves, the organization retains files and loses local judgment: which metric lies, which exception is safe, which workaround is temporary, what failed and who can authorize the next move.

The relay can extract a bounded operating corpus, verify integrity, test the successor against hidden cases, then keep a delta channel alive during transition. This is more defensible as a managed transfer service than a generic “AI memory” product.

Measure success as time to independent competent performance, retrieval accuracy, boundary correctness, rework and senior-attention minutes—not files migrated.

7. Private R&D and technical clean-room collaboration

Labs and companies can ask each other's situated agents to run approved analyses against proprietary datasets and return aggregates, falsifiers and provenance without moving the raw data. The value can be enormous when one cross-dataset result changes a drug, hardware, fraud, climate or network decision.

The option value is high; the sales cycle, domain validation and liability burden are higher. This is a later partnership direction, not an initial product.

8. Personal agent portability

People will change models, vendors and devices. A real transfer carries selected memory, current projects, exclusions, relationships, authority and tests. The receiving agent proves what it can retrieve and where it fails.

The market can be large, but consumer willingness to pay is unknown and model/platform vendors are structurally positioned to absorb portability features. Treat it as a long-term standards or premium concierge option, not the first business.

What the new research kills or downgrades

Do not build a generic relay platform

Microsoft Entra Agent ID already provides agent identities, owners, sponsors, lifecycle governance, access packages, conditional access and agent authorization. Google now exposes agent registries, gateways, traffic relationships, identity policies, semantic governance and audit trails. Identity and governance are becoming platform features.

Sources: Microsoft Entra Agent ID · Google Gemini Enterprise Agent Platform governance

Do not build the generic legal or payment layer

The American Arbitration Association and a coalition including Google, IBM, Circle, Wayfair and UiPath launched the open Legal Context Protocol in June 2026. It is designed to carry verifiable legal terms, consent, jurisdiction and recourse alongside agent transactions and to complement A2A, identity and payment protocols.

That makes legal context a standard to integrate, not a blank layer for Kairos to own.

Source: AAA Legal Context Protocol launch

Do not sell competence receipts as a standalone certification company yet

Certification frameworks and providers are already emerging. The Kairos competence receipt is still valuable as an internal acceptance instrument and a feature of a vertical service. It is not presently a defensible standalone business without a proven external benchmark, independent probe authorship, repeated buyer use and a recognized liability model.

Do not build a public consumer allowance or generic task market

The existing research already closes this branch: provider terms are fragile or prohibitive, exact credit marketplaces exist, task and agent markets have abundant supply but thin demonstrated paid demand, and generic payment/escrow infrastructure is available. Expiring capacity may lower execution cost. It is not the product or moat.

Do not use frontier models as pollers, schedulers or acknowledgement clerks

Zero-token sensors and code should own polling, hashing, dedupe, leases, retries, acknowledgement, archival and scheduling. The model wakes only for unresolved semantic work. The current relay's exhausted daily budget and large parked backlog are evidence that this distinction is economically load-bearing.

Recommended sequence

Phase 0 — repair the control plane

Before adding use cases:

  1. prove idle state causes zero model calls;
  2. install task-sensitive routing and actual-model/usage receipts;
  3. make MAX_ATTEMPTS apply to quota failures and replace red case 8's “will retry next run forever” expectation with bounded parking;
  4. separate deterministic ack/supersession from semantic judgment;
  5. establish one work ID, one lease, one retry budget and one tested human-visible failure path;
  6. clear or explicitly retire the existing backlog before interpreting new throughput.

Phase 1 — prove value inside MN/Kairos

Run two bounded pilots:

  1. payment exception replay against the measured renewal pool;
  2. paid-test conversion evidence bench across the July cohort.

For both, compare relay versus single-agent baseline on:

Proceed only if the relay changes an outcome or cuts human attention enough to repay its boundary cost.

Phase 2 — package one managed vertical service

Best first external wedge: a narrow verified-exception service for one commercial workflow where two parties already exchange evidence and disputes are expensive. Candidate order:

  1. billing/partner-charge exceptions in a sector where Kairos has access and a design partner;
  2. software/professional-service change assurance if no transaction-data partner is available;
  3. third-party incident tabletop and preparedness if a security partner is available.

Sell a result with an acceptance contract. Integrate existing identity, legal and payment standards. Do not ask buyers to adopt “the Kairos relay” as a platform.

Phase 3 — productize only repeated economics

Build shared product infrastructure only after at least one vertical produces repeated accepted outcomes and shows which components recur. The likely durable assets are:

Transport, general identity, generic orchestration and payment rails should remain replaceable dependencies.

Final calls

  1. Strongest MN/Kairos use: cross-boundary opportunity engine, proven first through payment recovery and paid-test conversion.
  2. Strongest compounding internal use: Standard of Performance ratchets across vital operations.
  3. Strongest protective internal use: incident rooms plus context/competence continuity.
  4. Best first external paid wedge: managed cross-company exception resolution in one narrow vertical.
  5. Largest external option: healthcare administrative exceptions and third-party incident coordination, both gated by reliability and liability.
  6. Best near-term proof of commercial craftsmanship: verified outcome packets for professional services/software changes.
  7. Wrong businesses: generic relay SaaS, generic agent marketplace, consumer allowance resale, identity/governance infrastructure, payment/legal protocol, or standalone competence certification.

Mammoth and authority boundary

Research, architecture, synthetic pilots and internal reversible validation can continue. A potentially material MN-linked product, external pilot, partner outreach, public release, spend, transfer or commercial pursuit must be registered and receive a signed Opportunity Schedule before the material step. Base service fees remain separate from upside participation.

Related vault sources

Kairos — Epistemic Federation and Sovereign Operators (2026-08-20) · Kairos — Agent Onboarding War Room Post (2026-08-19) · Kairos — Agent Onboarding and Parity Harness (2026-08-19) · Kairos — Expiring Agent Capacity Market Research (2026-08-20) · Kairos — Autonomous Agent Lanes (2026-08-12) · Kairos — Opportunity Hypotheses (2026-08-11) · Kairos — Payment Funnel Decision (2026-08-12) · Kairos — All-Hands Longitudinal Synthesis (2021–2024) · Kairos — Standard of Performance (2026-08-14) · Kairos — Mammoth Protocol Architecture (2026-07-24) · Mammoth Protocol v0.1 — discussion draft · 2026-08-20 — Autonomous runtime cost, safety, and ownership audit · ADR-063-agent-relay-unattended-pump-for-the-message-lanes

Browser rendering of the War Room source library. The vault remains canonical.