Skip to main content
AI, LuminexDoc, PDPA

Is It Safe to Send Company Documents to AI? What the Vendor Policies and PDPA Actually Say

6 September 2026 WinnerSoft Team
Is It Safe to Send Company Documents to AI? What the Vendor Policies and PDPA Actually Say

Whether it is safe to send company documents to AI comes down to three questions you can answer with documents rather than instinct — where the data is processed, how long it is kept, and whether it is used to train the model. The major AI providers publish written answers to all three, and those answers differ sharply between free tiers and commercial channels.

This article collects what the providers state in their own policies, what Thai law requires once a document contains personal data, and the questions a vendor should answer in writing before a project starts.

Are the documents we send used to train the model?

By default, no — provided you use a commercial channel. The three providers most Thai enterprises rely on say the same thing in their own documentation. But "by default" carries conditions, and the conditions are the part worth reading.

  • OpenAI — data sent to the API is not used to train or improve its models unless the customer explicitly opts in to share it.
  • Anthropic — by default it does not use inputs or outputs from its commercial products to train models; the exception is data submitted through feedback or with explicit consent.
  • Google Gemini API — only the paid services state that prompts and responses are not used to improve Google products. On the unpaid services, submitted content can be used to develop Google products and services.

That last point is where teams most often slip, and rarely at contract signing. It happens during evaluation — someone runs real documents through a free tier to see whether the technology works at all, long before legal is involved.

The difference between the two tiers is not model capability — it is what the provider is permitted to do with your data.

If they don't train on it, why keep logs at all?

To detect misuse, which is a separate matter from training. OpenAI states that abuse-monitoring logs are retained for up to 30 days unless longer retention is required by law, and offers Zero Data Retention controls that exclude customer content from those logs — subject to prior approval. Google states that for paid services it logs prompts and responses for a limited time solely to detect violations of its prohibited use policy and to maintain service security.

So the right question is not "do you store our data" — the answer is yes. It is "for how long, who can reach it, and can it be excluded?"

What does PDPA require once a document contains personal data?

Tax invoices, billing notes, contracts and official correspondence all carry names, identification numbers, addresses or signatures. When those documents enter an AI system, your organisation remains the data controller and the provider acts as a data processor operating on your instructions. Responsibility does not travel with the file.

When processing happens outside Thailand, the Personal Data Protection Act B.E. 2562 (2019) governs it through Sections 28 and 29, and the Personal Data Protection Committee issued criteria under Section 29 that took effect on 21 February 2025. An organisation using an overseas AI service should be able to name the legal basis it relies on for the transfer, and produce the processing agreement that supports it.

Three deployment shapes, and what each one costs you

Deployment Documents leave the organisation Used for model training What PDPA still requires
On-premise No No Retention policy, data-subject rights, and internal access control still apply
Cloud OCR API Yes, to the selected region Per the provider's contract Data processing agreement and a lawful basis for cross-border transfer
Commercial LLM API Yes No, by default policy The above, plus written confirmation that training is disabled and a stated log-retention period
Free LLM tier Yes Permitted under the terms Not appropriate for documents containing personal data
Security is decided by the channel you choose and the paperwork behind it — not by whether AI is involved.

On-premise is not the only definition of safe, and it is not free of cost. Someone has to run the machines and update the models, and accuracy on documents with no fixed layout is usually lower. What actually governs the risk is the paperwork, not the location of the server.

Human review is not only about accuracy — it is your audit trail

A system that routes only low-confidence fields to a person leaves a record of who changed which field, when, and from what value to what value. That record is exactly what internal audit and compliance ask to see, and it is something a manual keying process almost never produces.

Six questions to get answered in writing

  • In which countries are our documents processed, and which model providers are involved?
  • Is our data used to train or improve models, and which clause of the contract says so?
  • How long are logs retained, who can access them, and can our content be excluded?
  • Is there a data processing agreement to sign, and what lawful basis covers the cross-border transfer?
  • Does the system record who changed which field and when, and can that record be exported for review?
  • On termination, when are the extracted data and source documents deleted, and how is deletion confirmed?

These six extend the pre-signature checklist we published in "7 Questions to Ask a Document AI Vendor Before You Sign". That set covers accuracy and downstream integration; this one covers data and law.

How LuminexDoc answers them

LuminexDoc cross-checks every field with three AI models, and only the fields they disagree on or flag as uncertain reach a person. It requires no template per document layout and handles documents in more than 100 languages. It runs in daily production today across five Thai statutory document types: tax invoices, billing notes, receipts, debit notes and credit notes.

As for the six questions above, we answer them in writing before a POC starts rather than verbally in a meeting — because your compliance team will ask them again regardless.

Frequently Asked Questions

Does sending company documents to AI violate PDPA?

Not by itself. Your organisation remains the data controller, so you still need a lawful basis for processing, an agreement with the provider acting as processor, and — where processing happens abroad — compliance with Sections 28 and 29 of the PDPA B.E. 2562, including the Section 29 criteria effective 21 February 2025.

Do AI providers train their models on our documents?

On commercial channels, all three state they do not by default. OpenAI says API data is not used to train its models unless the customer opts in to share it. Anthropic says the same for its commercial products. Google draws the line explicitly at paid services; its unpaid services do use submitted content.

How long is the data retained?

It varies by provider. OpenAI states abuse-monitoring logs are kept for up to 30 days unless law requires longer, with Zero Data Retention available subject to prior approval. Google states that for paid services it retains prompts and responses for a limited time solely for policy-violation detection and service security.

Do we have to deploy on-premise to be safe?

No. On-premise removes the cross-border transfer question, but it adds infrastructure ownership and usually gives lower accuracy on documents with no fixed layout. What reduces real risk is a processing agreement, an explicit log-retention period, and an exportable audit trail.

Can we test with real documents without taking on risk?

Yes, if the frame is set first. Choose a document set with unnecessary personal data redacted, sign the processing agreement before the first file is sent, fix a deletion date for when testing ends, and never run real documents through a provider's free tier.