Data leakage
Every interaction transmits prompt content to the provider, where it may be logged, retained, evaluated, or used for training. On arrival it lands on multi-tenant infrastructure the customer cannot inspect.
Why the next phase of enterprise AI belongs to the firms that own their data and capabilities.
At first glance, enterprise AI adoption appears to be a success story.
Copilot drafting emails and integrated with Microsoft Teams, ChatGPT providing information about any topic, and maybe even a pilot project with a specialised vendor in a specific business unit. And yet, a quieter pattern has taken shape alongside the headline figures: AI has stalled at the edge of the organisation and has only rarely reached the core. High-stakes decisions, confidential analysis, strategic documents, and proprietary workflows have largely stayed outside AI systems — since routing that material through an external endpoint felt, correctly, like a bad idea.
The firms that will define the next decade of enterprise AI are already asking a different question from the one the market is answering. AI adoption is not about speed, but about building capabilities: this is the gap that will define the next round of competitive advantage. The cloud-AI model that allowed quick adoption is the same one that will make building capabilities impossible — the data leaves the firm, the engine belongs to someone else, and the legal regime governing both is set by a jurisdiction the firm does not vote in.
Thomas Kuhn’s argument in The Structure of Scientific Revolutions is useful here. Paradigms do not fall to critique; they fall to accumulated anomalies that the existing framework cannot absorb. The cloud-AI model has done its work and is now generating exactly that kind of anomaly. A paradigm shift is underway.
The cloud-AI architecture is familiar: prompts travel over the public internet to a third-party-hosted model; responses return; the provider charges per token. For experimentation, the arrangement is close to ideal. For sustained use of confidential data, it carries three costs that are easy to miss at adoption and expensive to unwind at scale.
Every interaction transmits prompt content to the provider, where it may be logged, retained, evaluated, or used for training. On arrival it lands on multi-tenant infrastructure the customer cannot inspect.
When a solution depends on a third-party engine, the engine’s provider captures the solution’s value over time — and gains a live view of how the firm operates.
The CLOUD Act, GDPR, the EU AI Act, and DORA all reach into externally hosted stacks — and are materially harder to evidence when model, runtime, and data pipeline sit outside the firm.
Every interaction with a third-party endpoint transmits prompt content to the provider. Depending on tier and contract, that content may be logged, retained for abuse monitoring, used for evaluation, or used for training. Training has a second-order consequence worth naming plainly: large models memorise portions of what they are trained on, and those memorised materials can be retrieved (Nasr et al., 2023). Agentic workflows compound the exposure — every tool call, every reasoning trace, every retrieval is another transaction with the provider.
Transmission is only the first layer. On arrival, prompts land on multi-tenant infrastructure — hardware and memory are typically shared, at the physical level, with dozens to hundreds of other customers depending on the service tier. Logical isolation between tenants is strong but not absolute, and the customer cannot inspect the boundary, audit who shares the silicon, or verify that the isolation holds. Shared infrastructure is a design choice that makes cloud economics work; it is also a design in which the leakage perimeter ceases to be contractual and becomes physical.
Enterprise contracts and zero-data-retention modes reduce this exposure, but do not eliminate it. The surface area is broader than the primary provider’s systems. In November 2025, OpenAI disclosed that names, email addresses, organisation IDs, and other metadata from API platform users had been exfiltrated — not from OpenAI’s own systems but from Mixpanel, a third-party analytics provider OpenAI had integrated on the front end. The attack vector was an SMS phishing campaign targeting a Mixpanel employee. The customer had no commercial relationship with Mixpanel, no way to audit its controls, and no reason to know it existed until the disclosure. Every link in the chain is an implicit trust assumption, and the chain is only as strong as its weakest link.
The key point is the asymmetry in consequence. The provider does not bear the downside of leakage the way the customer does. Nassim Taleb’s argument applies directly: different skin in the game produces different risk tolerances.
When a solution depends on a third-party engine, the engine’s provider captures the solution’s value over time. Sangeet Paul Choudary’s book Reshuffle (2025) traces the mechanism: as the engine becomes central to a solution, the engine provider acquires leverage across four compounding channels.
Pricing and terms flex in the provider’s favour.
The tool improves faster than anything built on top of it.
The engine evolves faster than dependents can adapt.
Widespread adoption of the engine erodes whatever differentiation the dependent firm had to begin with.
The solution provider does not capture part of the value added; instead, it misses out on learning and outsources its core capabilities to the AI vendor. The pattern is already playing out. In late January 2026, Anthropic released a set of plugins for its agentic Cowork platform — including one for legal workflows — competing with traditional SaaS suppliers. Days later, legal information, financial services, and enterprise software stocks sold off hard. Traders named the downturn the “SaaSpocalypse.” A horizontal AI provider took a vertical step into workflows where specialised vendors, some of them building on Anthropic’s own models, had raised substantial capital — and no one knew how to value those vendors anymore.
More broadly, a firm that routes its core knowledge work through a third-party engine hands that engine a live view of how the firm operates: which workflows matter, which decisions repeat, which expertise gets applied where. The dependency is not only commercial; it is informational, and it is the material from which the provider decides where to build next. Worse, because every competitor on the same engine gets the same capability, the tool never becomes an advantage. It becomes parity. The firm runs harder to stay in place, while the provider, watching the race, decides which race to enter.
The owner of a critical upstream input can always move downstream; the firm renting it cannot credibly move upstream.
This is not a new dynamic. It is a very old one, arriving faster than the vendors in its path expected. Schumpeter named the general form in 1942: creative destruction. Capitalist economies do not grow through steady accumulation; they grow through cycles in which new entrants with new capabilities render existing arrangements obsolete. Netflix understood the specific version earlier than most: streaming depended on studio licensing, and the studios — the upstream engine — could at any moment reprice or withdraw the content that made Netflix valuable. The response was to produce original content, so the solution’s performance no longer depended on an input the firm did not control. Enterprise AI is the same pattern on a shorter clock: the firm that builds its workflows, its institutional knowledge, and its differentiation on top of an engine it does not control is building on someone else’s roadmap.
The U.S. CLOUD Act requires American providers to produce customer data in response to a valid U.S. legal process, regardless of where the servers physically sit. Selecting a European region on a U.S.-headquartered hyperscaler does not remove that obligation. For any firm subject to the GDPR, the transfer of personal data to infrastructure under U.S. jurisdiction, however locally hosted, remains a live compliance question — and the European Data Protection Board’s 2024 opinion (EDPB Opinion 28/2024) sharpened the position by bringing the models themselves into scope.
Regulation adds a second layer on top of jurisdiction. The EU AI Act’s high-risk-system obligations become enforceable in August 2026. The act foresees penalties of up to €35 million or 7% of global turnover, and requires evidence — auditable logging, documented data governance, human oversight — that is materially harder to produce when the model, runtime, and data pipeline all sit within an external stack. “Sovereign cloud” offerings, whatever that means, address parts of the problem but do not resolve all of it.
For financial institutions, a third layer is already live. The EU Digital Operational Resilience Act (DORA), in force since January 2025, imposes direct obligations on ICT third-party risk: a full register of every critical provider and subprocessor, contractual rights of audit and access, documented exit strategies, and supervisory oversight of providers deemed critical to the sector. An AI stack built on a hyperscaler model, routed through an orchestration vendor, grounded on a managed vector database, and observed through a third-party analytics service is four ICT providers the institution must register, audit, and be able to exit. The architecture that makes cloud AI convenient is the same architecture DORA makes expensive to evidence.
Running AI workloads on owned infrastructure was substantially harder a couple of years ago. Three conditions have changed.
For fifty years, Moore’s Law has been the engine of enterprise technology economics: computing gets faster and cheaper on a predictable curve. AI hardware is on the same curve, and the generation-to-generation numbers inside a single product line make the point cleanly. VRAM is the binding variable for running modern AI models — a model either fits on a card, or across a small set of cards, or it does not.
A 70-billion-parameter model that required a data-centre footprint in late 2022 now runs on a four-card Blackwell workstation — with headroom for concurrent users, longer context, or several specialist models side by side. The hardware is no longer the binding constraint. Workloads that were not possible have become feasible; workloads that were possible have become cheap.
Two years ago, proprietary models held a clear lead over open-weight alternatives on essentially every benchmark. That lead still exists at the edge. But the binding question for the enterprise is not whether owned systems match the frontier on the tasks that make headlines — it is whether they match it on the tasks the firm actually depends on: extraction, classification, retrieval-augmented generation, drafting against proprietary context, and routine agentic flows. For those, the gap has closed. Six labs — Meta, Alibaba, DeepSeek, Zhipu AI, Mistral, and OpenAI — have shipped open-weight releases spanning reasoning, coding, and agentic tasks. Fine-tuning a domain specialist on an open base is a small fraction of the aggregate API spend required to serve the same workload externally at scale.
Clayton Christensen’s argument in The Innovator’s Dilemma is that market-leading incumbents systematically fail to lead disruptive shifts — not due to bad management, but the opposite. Rational investment leads toward sustaining innovations that serve the best customers on the metrics they already value, and away from disruptions that initially look lower-margin. Hyperscaler cloud AI sits squarely inside the pattern. Their best customers are cloud customers; their revenue depends on workloads running in hyperscaler data centres on hyperscaler terms. Pure on-premise, customer-owned infrastructure runs counter to that margin structure. The incumbents are not incapable of serving this market — they are structurally disincentivised from serving it well. That is the opening.
In 1973, the evolutionary biologist Leigh Van Valen proposed that species in a shared ecosystem run to stay in place: any advantage one acquires, another evolves to counter. He called it the Red Queen hypothesis, after Carroll’s line in Through the Looking-Glass — “it takes all the running you can do, to keep in the same place.” The idea describes enterprise AI with precision.
If every competitor can access the same broadly capable models, any improvement in the prompt library, the agent scaffold, or the workflow wrapper is immediately available to everyone operating on the same engine. Running is universal; so is standing still. Peter Thiel’s blunter version of the same dynamic is that “competition is for losers.” In commodity layers, participants exhaust each other without any of them building a durable advantage.
Balaji Srinivasan made the point in its most operational form: if everyone is using the same AI, then the competitive edge has to come from what isn’t in the AI. The moat moves from the model to everything around it — proprietary data, institutional decision-making history, client relationships, accumulated domain expertise — and the specific ways the firm’s people have learned to solve its problems. It is the unique way in which a firm organises and executes that turns a generic model into a specific capability.
The challenge is that the default cloud-AI architecture routes to convergence through the same channel that erodes idiosyncrasy. Every prompt, every retrieval-augmented query, every agent trace, every fine-tuning run is a transaction in which proprietary context is traded for a generic output. Enterprise contracts constrain how the engine’s provider uses that material. They do not change the fact that it leaves the firm’s perimeter. The strategic question over the next decade is therefore not how fast we can adopt AI. It is the extent to which we actually own our AI capability.
The moat is not necessarily abstract. In concrete terms, it is a system that becomes progressively better at operating within the firm’s context: its terminology, documents, decision patterns, and preferred outputs. Two mechanisms produce that compounding, and both require infrastructure that the firm controls.
Retrieval-based memory. Internal materials — reports, contracts, policies, meeting records, and manuals — are embedded and indexed so the model’s responses are based on the firm’s own facts and history rather than on generic training data. The model answers as if it has read the firm’s files, because it has.
Local adaptation. On owned infrastructure, interaction data — prompts, outputs, user corrections, accept/reject signals — refines the model for the organisation that produced it. Over time, the system learns which internal sources are authoritative, which answer styles the firm prefers, how its workflows route, and how its users correct the model when it is wrong. The feedback loop closes inside the firm’s perimeter.
The result is an AI capability that compounds. Every interaction becomes training data for the specific firm that produced it — a specificity that a competitor with access to the same base model but not the same history cannot replicate. This is what “your AI” means in the literal sense: a system whose weights and retrieval index reflect the firm’s actual operating knowledge, and whose every interaction sharpens both without leaving the firm’s perimeter. The dynamic-capabilities literature calls this a higher-order capability — not a fixed stock of resources but a learned ability to sense, seize, and reconfigure them as the environment shifts.
Owning an AI capability rather than renting one comes down to three things. The three are multiplicative: a firm that controls two of them and not the third has not achieved sovereignty. It has achieved the illusion of it. AI capability in the enterprise is a coordination problem, not an intelligence problem.
Of the knowledge the system learns from. Proprietary data, documents, decisions, and workflows are the assets. The question is whether they remain inside the firm’s perimeter during every stage of the AI lifecycle — training, fine-tuning, retrieval, inference — or are continuously transmitted outside it. When ownership is preserved, the firm’s knowledge compounds into its own capability. When it is not, it compounds into someone else’s model.
Of the infrastructure the system runs on. Who controls the weights, the runtime, the inference pipeline, the logs? Owned infrastructure means the firm can inspect, audit, modify, and port its AI capability. Rented infrastructure means dependency on someone else’s roadmap, pricing, and strategic priorities. The difference is the difference between a capability and a subscription.
Over who and what can use the system. This is the pillar most frequently under-attended, and the one whose absence quietly defeats the other two. A firm that owns its data and runs its model in a private cluster but allows unrestricted access to every employee, every connected SaaS tool, and every outbound integration has given up the governance the first two pillars were meant to provide.
The cloud-versus-on-prem question is not binary. It is a per-workload allocation. The table below is illustrative, not exhaustive — the specific mappings differ by industry — but the structure remains the same.
Non-sensitive inputs. Bursty or experimental. Commodity capability is sufficient. Examples: public-facing marketing copy; code scaffolding with open-source libraries; summarising public documents; meeting transcription for non-sensitive calls; prototypes and early-stage experimentation.
Consumption-priced public APIs. The flexibility and price/performance of hyperscaler-hosted models dominate.
Non-sensitive but steady-state, high-volume workloads where unit economics favour a reserved deployment over pay-per-token. Examples: customer-service chatbots on non-confidential content; public-facing knowledge retrieval at scale; production agent flows on non-strategic data.
Enterprise contracts on hyperscaler or vendor infrastructure — reserved capacity, VPC or private-link deployment, zero data retention and no-training terms, regional data residency.
The firm’s confidential and strategic assets — the material that constitutes the moat. Examples: proprietary research; M&A analysis on target-firm data; client-matter legal review; core trading logic; institutional decision history and knowledge retrieval; regulated data under EU AI Act high-risk, defence, certain healthcare and finance.
On-premise, owned hardware. Open-weight or internally fine-tuned models. Infrastructure and access fully under the firm’s control. No third party can query the system; no data leaves the perimeter.
The common mistake of the past three years has been running Tier 1 and Tier 3 workloads through the same architecture by default — almost always the cloud-AI default. The correction is not to invert the default. It is to recognise the tiers, audit the portfolio against them, and move misplaced workloads.
The shift isn’t eliminating cloud AI. It is that the strategic workloads — the ones that constitute the firm’s moat — require an architecture built for control rather than access.
Enterprise AI has had its first phase. The cloud made experimentation cheap and adoption fast — exactly what the moment required. The moment has changed. The workloads that will define competitive advantage over the next decade are those the cloud-AI default cannot safely handle.
Our conviction is that on-prem deployments will move from edge case to norm over the next several years for the workloads that matter: because the hardware keeps improving, because open-weight systems match proprietary systems on the tasks enterprises actually depend on, and because firms will ultimately want control over their most valuable asset — the accumulated knowledge that no external model can provide on their behalf. The argument is not unconditional. If the frontier re-accelerates beyond the reach of open weights, or if hyperscaler controls resolve the three costs in ways they have not so far, the calculation shifts. We believe neither is likely.
The reader who disagrees will find those the variables to watch. Paradigms do not end — they are replaced. Kuhn’s point and Van Valen’s are the same in different vocabularies: the running never stops. The firms that adopt AI as a tool will be overtaken by those that adopt it as a capability — one they own, govern, and compound within their own perimeter.
Who owns the knowledge that trains your system?
Who owns the infrastructure that runs it?
Who decides who gets access?
These are strategy questions, not procurement questions. They determine whether AI capability meets Barney’s (1991) test for sustained competitive advantage — valuable, rare, imperfectly imitable, and non-substitutable — or becomes a continuous export of the differentiation that used to be yours.
Cloud AI is a fight for survival. Your AI is a fight for victory.
Myrstack is live on the ground today, operated and owned by its users.
Request Myrstack