BACK TO ALL POSTS
tools

Does Your Martech Train AI on Your Data? 54 Vendors, 6 Clear Answers

Dev Anand

AI Tool Reviews & Automation Safety · 2026-08-16 · 11 min read

Part of: Research
Does Your Martech Train AI on Your Data? 54 Vendors, 6 Clear Answers

Key Takeaways

  • Of 54 martech vendors, 14 (25.9%) state that customer data does not train their own AI models. Only 6 of those say it in a legal document.
  • 14 vendors (25.9%) reserve the right to train their models on customer data in their terms or privacy policy. Four offer an opt-out.
  • 16 vendors (29.6%) have no findable statement either way, in legal documents or anywhere else on their site.
  • 6 vendors address only their third-party AI providers, and 4 have nothing beyond the Google Workspace API sentence.
  • Vidyard's privacy policy goes furthest: it may license customer videos to third parties for AI training, with an opt-out for free-plan users.

Does Your Martech Train AI on Your Data? 54 Vendors, 6 Clear Answers

By Dev Anand, AI Tooling & Automation Safety. Last updated: 2026-08-16

Every martech vendor now ships AI features. Almost every buyer now asks the same question in procurement: does my data train your models? We wanted to know how many vendors answer that question in public, in writing, before anyone has to ask.

The honest summary is that most do not, and the ones that do are not the ones you would guess.

How many martech vendors say whether they train on customer data?

We started with the same 54 live vendors as our pricing transparency audit, spanning CRM, email automation, sales engagement, data enrichment, SEO and content, analytics, AI writing, social tooling, video and AI SDR platforms. For each, we crawled the privacy policy, terms of service and every page reachable from the legal, trust and security sections, plus a fixed list of common AI-terms paths. Then, for any vendor whose legal pages said nothing, we ran a site-restricted search and read what came back.

Every vendor landed in one of five buckets.

What the vendor says Vendors Share
Commits not to train its own models on customer data 14 25.9%
Reserves the right to train its own models on customer data 14 25.9%
Addresses only third-party AI providers 6 11.1%
Only the Google Workspace API boilerplate 4 7.4%
No statement found anywhere 16 29.6%

The two largest groups are exactly the same size, and they point in opposite directions. A quarter of the corpus promises not to train. A quarter reserves the right to. The remaining half either dodges the question or never gets asked it in a document.

Which vendors reserve the right to train on your data?

This is the group buyers most need to know about, and it is not obscure. All 14 make the statement in a binding document, because a right that is not in the contract does not exist.

Vendor What the document says Opt-out?
HubSpot "We may also use Customer Data to train our AI models" Yes, in settings
Mailchimp May use "Inputs and Outputs, including Customer Data, for machine learning purposes in order to develop and improve the AI Model" No
ActiveCampaign May "internally use Marketing Content to help us train and improve the Services" No
Apollo "may use Customer Data... as inputs to train these internal models" No
Ahrefs "You agree that we may use User Content... to train our machine learning models" No
Hootsuite May use "Customer Content and Outputs to develop and improve... our machine-learning technologies", not shared with other customers No
Sprout Social "may use the AI Content to train, develop, and improve the AI Features" No
PostHog "may use data you submit... to train our own internal models" Yes, forward-looking only
Jasper Will "cease using your Customer Property for model training and improvement" on verified request Yes, on request
Loom (Atlassian) Uses information "for development, training, or fine-tuning of machine learning and artificial intelligence models" No
Vidyard May "test, train and improve AI technologies and models", and may license videos to third parties for AI training Free plan, in settings
Artisan Processes personal information "to train, fine-tune, and evaluate our language models" No
Lusha Will not train "public AI", but AI features "may be trained in Lusha's local and offline environment... mainly with regard to Customers' metadata" No
Pipedrive Beta AI features may use "Client Data, including your Input and Output to train or refine such capabilities" for that customer's beta service No

Two things stand out. First, HubSpot, PostHog and Jasper are the only three that offer any opt-out, and PostHog's terms are candid that opting out "can't undo any model training that's already taken place." Second, Vidyard is the only vendor in the corpus whose privacy policy contemplates handing customer content to third parties for their AI training, a materially different thing from improving its own product.

None of this is hidden. It is in the documents. It is just not in the documents most buyers read.

Want to put this into practice?

Reachium automates LinkedIn outreach, content publishing, and inbox management in one platform.

Start Free →

Which vendors promise not to train, and where do they promise it?

Fourteen vendors make the commitment. Where they make it matters more than the fact that they make it.

In a legal document (6): Semrush's AI Services Terms say it "does not train or develop its models, and does not allow its third-party providers to train or develop its models on User Input." Writesonic's privacy policy says it does "not use Customer Data to train or fine-tune any general-purpose, foundation, or large-language model offered by Writesonic or by any Model Provider." 11x's terms say it "shall not use any Customer Data to train any artificial intelligence or machine learning models." Frase, Mixpanel and Dreamdata make narrower but still contractual commitments: Frase excludes training "in a manner that would... apply those learnings broadly across other customers," Mixpanel permits training "solely for the benefit of an individual customer," and Dreamdata's AI measures page states "there is no pooled training set."

Somewhere else on the site (8): Zoho, Clay, Clearscope, Amplitude, HockeyStack, Copy.ai, Typeface and Relevance AI all say some version of "we don't train on your data," and all say it on a trust center, docs page, security page, product page or blog post rather than in their terms. Clay's is on its trust FAQ. Clearscope's is in a blog post. Copy.ai's is a headline on its security page. Zoho's appears on product pages for SalesIQ, Projects and Desk, but not in the privacy policy or terms we fetched.

We do not doubt any of these eight. But a promise on a marketing page can be edited without notice and is not what a procurement lawyer will accept in a redline. If you are one of these vendors, moving one sentence into your terms would move you from the second list to the first.

What does it mean when a vendor only mentions Google?

Four vendors, Copper, Outreach, Regie.ai and Qualified, have precisely one sentence about AI training in their legal documents, and it reads something like "data obtained through Google Workspace APIs is not used to develop, improve, or train generalized AI and/or ML models."

That sentence is not a policy choice. It is required by Google's API Services User Data Policy of any app that reads Gmail or Workspace data, and it covers only data that came through those APIs. A CRM record you typed in, a call transcript, a sequence you wrote: none of it is Workspace data, and none of it is covered.

For a buyer, the practical reading is that these four vendors have not addressed the question at all. The sentence is there because Google made them put it there.

Six more vendors, Salesforce, Klaviyo, Customer.io, Braze, Heap and Anyword, address only their providers. Klaviyo's terms say "Customer Data will not be used to train third-party foundation models," and Customer.io's say its providers "are contractually prohibited from using your inputs or outputs to train their foundation models." Both are real commitments. Neither says anything about the vendor's own models, which is where the question actually lives now that most vendors fine-tune something.

Which vendors say nothing at all?

Sixteen vendors have no findable statement: Close, Iterable, Salesloft, Lemlist, Instantly, Smartlead, ZoomInfo, Cognism, Seamless.ai, Surfer, MarketMuse, Taplio, Buffer, Demio, Livestorm and AiSDR.

We looked harder at these than at the others, because "not found" is a serious thing to publish. Each got the legal crawl and then a site-restricted search for the vendor's name with "train" and "customer data." Iterable has a dedicated Additional AI Terms of Use page that does not mention training. Instantly's terms contain a disclaimer that it "is not responsible for the Third-Party Services' handling of your Inputs or Outputs, including for use in their model training," which is a warning, not a commitment. Smartlead's only statement is a help-center note that its Slack app does not train on Slack data. Salesloft's AI page says "your data remains yours," which is an ownership claim rather than a training one.

Some of these vendors surely have an answer in a signed enterprise agreement or a gated trust portal. The measurement here is what a buyer can find before signing, and for these sixteen the answer is nothing.

Want to put this into practice?

Reachium automates LinkedIn outreach, content publishing, and inbox management in one platform.

Start Free →

What should a buyer do with this?

Three things follow directly from the data.

Ask the question in writing before you sign, and ask about the vendor's own models, not just its providers. Half the vendors that address the topic at all address only OpenAI and Anthropic.

Read the AI-specific terms page if one exists, and check the vendor's changelog for when the AI features actually shipped. In this corpus, the decisive sentence for HubSpot, Hootsuite, Sprout Social, Customer.io, Semrush and Mixpanel lived on a separate AI terms page, not in the main terms of service.

Do not treat a trust-center FAQ as a contract. Eight of the fourteen "we don't train" commitments are on pages the vendor can edit tomorrow. Ask for the sentence in the order form.

The broader pattern connects to something we see across the martech stack: the more tools you run, the more of these documents nobody on your team has read. It is one more argument for consolidating onto fewer platforms whose terms you have actually checked.

What are the limits of this audit?

Every statement was read on 2026-08-16 and legal pages change without notice. "Not found" means not found by our procedure: a crawl of legal, trust and security paths plus a site-restricted search. A statement could exist in a signed contract, a gated portal or a help article the search did not surface. The audit verifies what vendors say, not what they do. And a few bucket boundaries are judgment calls; Mixpanel and Pipedrive both reserve narrow customer-scoped training rights and landed on opposite sides of the line, Mixpanel because its right is per-customer only and Pipedrive because its right is a general beta-features clause. The full per-vendor table with every decisive sentence and URL is published alongside this piece so anyone can disagree with a specific call.

FAQ

How many martech vendors commit not to train AI on customer data?

14 of 54 in this audit make that statement somewhere on their site. Only 6 (Semrush, Writesonic, 11x, Frase, Mixpanel, Dreamdata) make it in a legal document such as terms of service, a privacy policy or an AI terms page.

Which major martech vendors reserve the right to train on customer data?

HubSpot, Mailchimp, ActiveCampaign, Apollo, Ahrefs, Hootsuite, Sprout Social, PostHog, Jasper, Loom, Vidyard, Artisan, Lusha and Pipedrive all state in their terms or privacy policy that they may use customer data or content to train or improve their models. HubSpot, PostHog and Jasper offer an opt-out.

Does HubSpot train AI on customer data?

HubSpot's privacy policy states it "may also use Customer Data to train our AI models" and describes an opt-out: "If you opt out, we will no longer collect Customer Data to train our AI models." Opting out does not remove access to HubSpot's AI features.

What does the Google Workspace API sentence mean?

It is a disclosure Google requires of any application that accesses Gmail or Workspace data, stating that such data is not used to train generalized AI models. It covers only data obtained through those APIs. Four vendors in this audit have no other statement about AI training.

Is a "we don't train on your data" statement on a trust page binding?

Not in the way a clause in the terms of service is. Eight of the fourteen commitments found here appear only on trust centers, docs, security pages or blog posts, which a vendor can change unilaterally. Buyers who need the commitment should ask for it in the contract.

Want to put this into practice?

Reachium automates LinkedIn outreach, content publishing, and inbox management in one platform.

Start Free →

Sources


GTMStack publishes original audits of the AI marketing stack. If you want the next one when it lands, the newsletter is the place to get it.

Want to automate what you just learned?

Reachium turns these strategies into automated LinkedIn campaigns that book meetings on autopilot.

Try Reachium Free

MORE FROM GTMSTACK