The short answer
Ask the vendor one question in writing: is the learning loop scoped to my account, or pooled across all customers? Cross-account pooling gives new accounts a warm start and surfaces patterns a single account cannot see. Single-tenant learning gives clean boundaries and simpler contracts. Neither is wrong. What matters is which one you bought, and whether the contract says so.
This is a buyer’s checklist, not an argument for one architecture. The failure mode is rarely picking the wrong model; it is that nobody asked, and the contract turned out to describe something other than what the deck implied.
“Does it learn?” and “whose data does it learn from?” are different questions, and vendors answer the first to avoid the second. We covered the first in do AI sales agents actually learn. This post is about the second: what goes into the vendor’s feedback layer, whose data sits beside yours, and what comes back out.
What actually gets pooled
“We use aggregated learnings across our customer base” covers three different things. Vendors are rarely lying; they are imprecise about which they mean, and the three have risk profiles that are not comparable.
| What is pooled | Concretely | Re-identification risk | Competitive-leakage risk | What to ask for |
|---|---|---|---|---|
| Raw content | Recordings, transcripts, message bodies, CRM notes, uploaded lists | High. Personal information by any definition, not de-identified by being copied to a shared store. | High. Scripts, objection handling and lead lists are readable. | A flat prohibition on use outside your instance and deletion terms. |
| Derived features and model artefacts | Per-lead scores, embeddings, classifiers, weights fitted partly on your data | Contested. Regulators do not treat “derived” as a synonym for “anonymous”. | Moderate to high, hard to audit. A model fitted on your outcomes encodes what works on your buyers. | Whether features are per-tenant or global, and whether global artefacts are fitted on your data. |
| Aggregate outcome statistics | Category benchmarks: contact windows, channel response rates by industry and country | Low, if aggregation is genuine and cohort sizes enforced. | Low. Closer to a market benchmark than to your data. | A minimum cohort size, so no one account dominates a statistic. |
Buyers who object to pooling mean row one. Vendors who defend it describe row three. The argument is about row two, where nobody has clean language.
The European Data Protection Board addressed row two in Opinion 28/2024, adopted December 2024: “AI models trained with personal data cannot, in all cases, be considered anonymous”, with anonymity assessed case by case. If a vendor’s answer to row two is the single word “anonymised”, that is a claim to be evidenced, not a category that ends the conversation.
The law you are actually relying on
Neither Australia nor the EU bans cross-account learning. Both have something more useful: data can only be used for the purpose it was collected for.
The Office of the Australian Information Commissioner’s APP 6 guidelines state that an APP entity holding personal information “can only use or disclose the information for a particular purpose for which it was collected (known as the ‘primary purpose’ of collection), unless an exception applies”. The commercially relevant exceptions are consent, and the individual reasonably expecting a related secondary use.
In the EU, Article 5(1)(b) of the GDPR requires personal data be “collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes”, per the Official Journal text on EUR-Lex. That is the clause a pooling arrangement has to survive.
The OAIC has addressed this product pattern directly. Its guidance on commercially available AI products notes that “Some commercial AI products will include terms or settings that allow the product owner to collect the data input by customers for further training and development of AI technologies”, warning of “the potential for personal information input into an AI product then surfacing in response to a prompt from another user”.
The consequence is uncomfortable. If your vendor uses prospect details you collected to improve a shared model serving other customers, you are the one who must justify that as a permitted secondary use. You are typically the controller, and a contract allocates work without moving obligations.
What “we don’t train on your data” does and does not cover
The phrase does more work than it can bear, because an AI sales stack has at least two vendors and it usually covers only one.
Anthropic’s Commercial Terms of Service, in the Customer Content section, state: “Anthropic may not train models on Customer Content from Services”, Customer Content being Inputs and Outputs. Its Data Processing Addendum goes further. Clause B.3.c commits that Anthropic will not, except as otherwise permitted by Applicable Data Protection Laws, “combine Customer Personal Data with personal data that Anthropic receives from or on behalf of another person or persons, or collects from its own interaction with the data subject”. Clause H.1 sets return on request, and deletion, within thirty days of termination or expiry, subject to stated carve-outs including legal retention requirements.
B.3.c is the actual anti-pooling clause, and it is rarer than the training clause. A vendor can honestly say it does not train on your data whilst still combining your data with another customer’s for analytics or benchmarking. Separate promises; ask for both.
Now the part buyers miss. That commitment binds the model provider, not the sales-agent vendor on top of it, and it is entirely compatible with that vendor pooling every transcript into its own feedback layer. A vendor who answers your pooling question by forwarding their model provider’s privacy page has answered a different question.
What the phrase routinely excludes
- The vendor’s own orchestration and analytics layer, where the pooling decision is actually made.
- Metadata and outcome labels: who answered, when, whether they booked. Rarely counted as “content”, often the most useful signal there is.
- Sub-processors, human review, and anything retained after termination. “We don’t train on it” is not “we delete it”.
What a DPA should say
- Roles named. Who is controller and who is processor, in words.
- Processing limited to documented instructions, with “improving the Services” excluded or defined tightly enough to be enforceable. An open-ended service-improvement clause is a pooling permission wearing a different hat.
- A combination clause in the shape of B.3.c above, or carve-outs stated explicitly.
- Sub-processor list and change notification, with a right to object.
- Deletion on termination within a fixed period, covering backups and derived artefacts, not just records.
- Transfer mechanism for cross-border processing. In the EU that is usually the Standard Contractual Clauses. In Australia, the OAIC’s APP 8 guidelines note that an entity disclosing personal information overseas “is accountable for any acts or practices of the overseas recipient in relation to the information that would breach the APPs (s 16C)”.
Our voice-AI vendor security questionnaire covers the wider procurement surface; this post is one question from it, examined in depth.
The competitive-leakage question
The scenario buyers worry about: my vendor serves several businesses in my vertical, possibly a direct competitor. If the system learns from all of us, am I paying to educate a machine that then works for them?
Sceptics understate how little of what a cross-account system learns is proprietary to you. That Tuesday mornings beat Friday afternoons in a given industry, that one opener loses to another, that a form filled after hours converts differently: these are properties of the market, not of your business. Advocates understate how much of it plainly is yours: your list sources, qualification thresholds, objection library and offer construction. Transporting those to a competitor is not benchmarking, it is transfer, and the boundary is drawn by the vendor’s engineering rather than your intuition.
So the useful question is not “do you pool?” but “what is the unit of learning?” A system that learns category-level patterns is structurally different from one that lifts your winning script into a competitor’s sequence. Ask which, and ask for the boundary in a sentence you could hold them to.
Where Zian sits
Applied to ourselves rather than around ourselves: Zian is on the cross-account side of this trade-off, and we have said so publicly rather than in the fine print. Our team has been running outbound acquisition since 2017, at one point operating campaigns for roughly a hundred businesses in the same vertical simultaneously, which is exactly the arrangement this section describes. Our learning engine tracks around 420,000 data points across more than 10,000 leads a day.
The argument for it is the one above: a pattern that takes one account two quarters to confirm is visible sooner when the sample is larger, and what we have said publicly is that the correlation can then be pushed back across accounts. We are not going to argue that the trade-off in this post applies to everyone except us.
What we have not published, and will not improvise in a blog post, is the contractual detail: what is retained and for how long, and where the boundary sits between a category-level pattern and anything specific to your campaigns. Those are the ten questions below, and they are ours to answer in writing like anyone else’s rather than ours to pre-answer here. Private model deployment on customer infrastructure is on the capability list, and we cover what it involves in AI data sovereignty in Australia — though by the argument made higher up, moving the model layer inside your own boundary is a different question from what any vendor’s own orchestration and feedback layer does, and should be asked separately. Zian is in waitlist beta, with no public pricing and no self-serve signup, so terms are settled with a person rather than on a page.
The questions to put to a vendor in writing
In writing, because a confident verbal answer is not a commitment and asking is itself the test.
- Is your learning loop scoped to my account, pooled across accounts, or both at different layers?
- Of raw content, derived features and aggregate statistics, which crosses the boundary?
- Does any model or artefact serving other customers get fitted on my data? If so, can it be removed, and how would you demonstrate that?
- If you do not train on my data, does that cover your own systems or only your model provider’s?
- Is there a clause preventing you combining my customer data with another customer’s? Point to it.
- Do you serve competitors of mine, and what prevents my material reaching their campaigns?
- What is your sub-processor list, and how am I notified of changes?
- On termination, what is deleted, within what period, and does that include derived artefacts?
- Is a single-tenant or private deployment available, and what changes if I take it?
- Which of those answers are in the contract today, and which would need adding?
Question ten separates a vendor who has thought about this from one who has not. Any salesperson can answer one through nine. Only a vendor with real internal clarity can say, without checking, which of those answers are contractual today.
How we sourced this
Every quotation above was checked by opening the primary source on 28 August 2026, not by reading a summary: paragraph 6.2 of the OAIC’s APP 6 guidelines, the OAIC’s guidance on commercially available AI products, Article 5(1)(b) of the GDPR as published in the Official Journal and served by EUR-Lex (reachable from our host and read directly), and the executive summary of the EDPB’s PDF of Opinion 28/2024. The contractual example is Anthropic’s Commercial Terms of Service (marked effective 17 June 2025 when checked) and its Data Processing Addendum (marked effective 24 February 2025), chosen because the pages were retrievable and the clauses unusually explicit, not as an endorsement.
What we could not verify: OpenAI’s enterprise privacy page returned an HTTP 403 to our host, so we left it out rather than quote it second-hand. Terms change, so open the current version before relying on it. Figures for Zian’s own history and scale are first-party operating data for which no external source exists.
Frequently asked questions
Is cross-account learning legal?
There is no blanket prohibition in Australian or EU law; what applies is purpose limitation. The OAIC’s APP 6 guidelines state that an entity “can only use or disclose the information for a particular purpose for which it was collected (known as the ‘primary purpose’ of collection), unless an exception applies”. So the question is whether the secondary use fits an exception such as consent or reasonable expectation. That analysis is yours, because you are usually the controller.
If a vendor says the pooled data is anonymised, is that the end of it?
No, it is the start of an evidence question. The EDPB’s Opinion 28/2024 concluded that “AI models trained with personal data cannot, in all cases, be considered anonymous”, with anonymity assessed case by case. Ask what technique was used, at what cohort size, and whether it was tested.
Does “we don’t train on your data” mean my data stays in my account?
Not necessarily. It covers model training, not data combination, analytics or benchmarking, and often refers only to the underlying model provider rather than the sales-agent vendor’s own systems. Ask for a separate commitment on combination, in the shape of Anthropic’s DPA clause B.3.c.
Is single-tenant learning simply safer?
It is simpler to contract and to explain, which has real value, but it is not free. A single-tenant system has only your history to reason from, so it waits longer for statistical significance and cannot see patterns your volume never produces. You are choosing between a cleaner boundary and a faster feedback loop.
What should I do if my vendor serves a direct competitor?
Ask what the unit of learning is. Category-level patterns about when buyers answer the phone are market properties and were never exclusively yours. Your scripts, list sources and offer construction are another matter. Get the boundary written down and ask what enforces it.
Apply for partnership
If data governance is a gating requirement rather than a checkbox, that conversation is worth having before a demo rather than after one. Zian is in waitlist beta and we take partners selectively. Apply For Partnership and bring the ten questions with you.