Vendor riskAITPRM Health™August 2026

Are Your Artificial Intelligence Vendors Training Models on Your Patient Data?

The permission is rarely in the business associate agreement. It sits in the services agreement alongside it.

Every week, healthcare organizations are pitched new tools promising to automate clinical documentation, speed up revenue cycle work, or optimize scheduling. The pitch deck displays a compliance badge prominently. Look past the marketing and into the fine print of the master services agreement, and a critical legal gap frequently emerges.

Many commercial providers include broad data usage rights in their standard terms. Unless procurement and legal know exactly where to look, an organization can inadvertently grant a vendor permission to use patient-derived data to train its commercial models.

The product improvement trap

Vendor contracts routinely contain clauses allowing the vendor to retain, analyze, and process customer content for system optimization, algorithm tuning, or product improvement.

That sounds unremarkable for traditional software. In the context of a generative model, it creates severe exposure. If the terms permit the model to learn from your clinical notes or patient interactions, protected health information or derivative signals may be absorbed into model weights, a process that is effectively impossible to reverse.

What the agreement says
The blind spot beside it
The vendor will safeguard protected health information under HIPAA
The services agreement separately allows data usage for product development or model evaluation
The vendor uses encryption in transit and at rest
Derivative data such as embeddings and summaries is rarely defined as customer-owned
The vendor executes a business associate agreement
Downstream foundation model providers are not bound by flow-down terms

What happens to the derivatives

When patient data is processed by a model, the system generates intermediate products: vector embeddings, summary tokens, and inferences. Does your agreement explicitly define who owns them? If the contract is silent, the vendor may claim ownership of those mathematical representations of your patient data, and use them to refine products sold to your competitors.

Three steps before signature

01
Enforce default-off training terms

Invert the language so that training, fine-tuning, and performance evaluation on customer content are strictly disabled by default.

02
Assign ownership of derivatives

Define all prompts, outputs, vector embeddings, and inferences as customer-owned records in the contract itself.

03
Inspect subprocessor flow-down

Require written documentation proving those protections extend to every third-party model host executing the underlying processing.

Key takeaways
A compliance badge is not enough. Training rights often sit in general services terms, not the agreement you reviewed.
Derivative data has to be owned. Embeddings and generated summaries must be defined contractually as your property.
Subprocessors must be verified. Protections have to cover the entire chain of model providers.
Close your vendor risk gap.

A framework-aligned assessment your own team runs, covering the domains a standard questionnaire does not reach.

Talk with our team
← All posts
Note

Published for general informational purposes. This material describes regulatory and operational practices and does not constitute legal advice, and it does not create an attorney-client relationship. Statutory requirements change, and their application depends on your organization’s facts. Consult qualified counsel regarding your obligations.

Next step

Bring us a vendor or a board question.

Thirty minutes with a compliance lead. Use your own categories, locations, and obligations, and if the fit is wrong for your scope we will tell you.