Is there data we simply should not send to an AI model?
Yes. Payment card numbers, bank details, government identifiers, passwords and API keys, alarm and gate codes, and anything covered by a regulatory regime you have not specifically built for. The practical test is not whether the model can handle it, but whether you would be comfortable seeing that string in an application log, a retry queue, an error report or a support ticket. If not, redact it before the call is made.
The log test
Anything you send to a model may also appear somewhere you did not intend. Requests get logged for debugging. Failed jobs sit in retry queues with their payloads intact. Exceptions ship stack traces to a monitoring service. A support ticket includes the request that failed.
So the question is never just "is the model provider trustworthy." It is "am I comfortable with this value existing in six places I do not think about." That reframing settles most arguments quickly.
The exclusion list
- Payment instruments. Card numbers, CVVs, bank routing and account numbers — including when spoken aloud in a recording.
- Government identifiers. Social security numbers, driver's license numbers, passport numbers.
- Secrets. Passwords, API keys, tokens. Obvious, and still pasted into chat tools weekly.
- Physical access codes. Alarm codes, gate codes, lockbox combinations from service notes.
- Regulated categories you have not scoped. Clinical detail, background check results, protected employment records.
- Whole databases. Sending a full customer export to answer a question about one account is the most common overshare there is.
Minimization is an architecture decision
Policy language telling employees not to send sensitive data helps, but the durable fix is structural: the system only assembles the fields the task requires. A call summarizer needs the transcript text, not the customer's payment record. A scoring pass needs the conversation, not the account balance.
When the prompt is built from an explicit field list rather than a whole record, exclusion becomes the default and inclusion becomes deliberate. That is a normal part of designing the integration layer, and it is far easier than auditing free-form usage after the fact.
Put redaction at the earliest possible point
Redact at the transcription output, not at the report. Every stage downstream of the redaction point is clean; every stage upstream is not. Moving that boundary one step earlier removes whole categories of copies from consideration.
Then verify. Run a periodic scan of your transcript store for the patterns that should never be there. Finding zero is reassuring. Finding some tells you exactly where the pipeline leaks — and the same scan belongs on anything feeding your reporting layer.
Topics: data minimization · redaction · security · PII
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.