After this module you can: decide in 5 seconds whether something may go into a public AI tool, use the approved tool instead, and rewrite a prompt so the structure stays and the specifics go.
✅ Fact-checked 2026-09-03. Vendor data-handling defaults are tier-dependent and change over time — everything below states which tier it applies to and is accurate as of that date. If you're reading this much later, the principle still holds; the vendor specifics may not.
1 · The P&L summary
Quarter-end. Someone needs an executive summary of the quarterly P&L for a deck due in an hour. A public AI chatbot writes beautiful summaries. Paste, prompt, done — the summary is genuinely great, and it took two minutes.
The problem is invisible. The full P&L — revenue by operator, margins, salary lines — is now sitting on someone else's servers, under someone else's terms, outside every contract Tom Horn has signed. We can't audit it, can't delete it with certainty, and can't promise a partner or regulator it's contained. This is the second of our two real incidents, and nobody involved meant any harm. The unsafe route was simply faster.
The question is never "is this AI tool evil?" It's: would I post this in a hotel lobby? A public chatbot on a personal account is a lobby — friendly, useful, and not ours.
2 · Where a prompt actually goes
When you hit Enter, the text doesn't stay in the chat window:
1
Your device → the vendor's servers. The full prompt travels (encrypted in transit) to the AI company's infrastructure. Nothing sensitive is removed on the way — whatever you pasted, they received.
2
Retention & logs. Most tools store your conversation history so you can revisit it — and vendors keep operational logs too. How long, and who can see them, depends on the product and tier.
3
Possible human review & training — tier-dependent. On consumer/personal tiers, conversations are commonly retained and may be used to improve models, or reviewed by humans, depending on the vendor and your current data-control setting (as of 2026-09-03). On business/enterprise/API tiers, major vendors state they do not train on customer inputs by default, and offer contractual controls, shorter or zero retention, and audit features.
The three big ones, as of 2026-09-03 (personal/consumer accounts)
ChatGPT (OpenAI), Free/Plus: conversations are used to improve models by default; you can switch it off under Settings → Data Controls → “Improve the model for everyone.” Team/Enterprise/API are not trained on by default. Deleted chats are removed from OpenAI systems within ~30 days.
Claude (Anthropic), Free/Pro/Max: since the 2025 consumer-terms update you make a choice; if you allow training, chats can be retained up to 5 years, otherwise the standard is ~30 days. Commercial/API/Enterprise: not trained on by default.
Gemini (Google), consumer app: with activity saving on, a sample of conversations can be reviewed by humans and kept up to 3 years, stored separately from your account — Google's own guidance says don't enter anything confidential you wouldn't want a reviewer to see. Google Workspace/Vertex enterprise terms are different: business data isn't used to train the general models.
These defaults are exactly the kind of thing vendors change — which is the real lesson. Don't memorise the table; use the approved tool, or sanitize.
4
Outside our contracts — the part that's true in every case. A paste into an unapproved or personal tool moves company data outside THG-approved contracts, controls and visibility. Even if the vendor behaves perfectly, we can no longer answer "where is this data and who can access it?" — and for a regulated supplier that question gets asked.
Why we're not saying "the AI trains on everything you type": because it isn't universally true, and advice you can catch out is advice you stop believing. The honest rule: which tool, which account, which tier decides what happens to a paste. An approved business-tier tool under our contract is fine. The same vendor's free consumer app on your personal login is not — that's "shadow AI," entirely outside our agreements.
3 · The NEVER list — what stays out of public AI tools
🎰 Player data & game results — regulated data; session logs, bet histories, IDs (even "just player IDs" — pseudonymised IDs still count as personal data).
💶 Financials — P&L, salaries, forecasts, pricing.
📜 Contracts & legal documents — drafts included.
🔑 Credentials — passwords, API keys, tokens. Never, anywhere, in any tool.
🎮 Unreleased game data — math models, RTP sheets, roadmaps, unannounced titles.
The 5-second test: is anything in my paste on this list? Yes → approved internal tool only, or sanitize first. Unsure → treat it as yes and ask.
The riskiest habit isn't a category on the list — it's the personal account. Work data in a chatbot logged into your private email is outside every THG contract, invisible to us, and mixed into your personal history. Use your work account in the approved tool, every time.
4 · Sanitize the prompt: structure stays, specifics go
The good news: AI doesn't need your real numbers to help you. It needs the shape of the problem.
❌ Before (real data leaves the building)
Summarise this P&L for the board:
Operator BetNordic: €412,300 rev, 61% margin
Operator LuckySpin: €188,450 rev, 44% margin
Salaries: €96,200 · Server costs: €31,900 …
✅ After (same help, nothing sensitive)
Write a 3-bullet executive summary template
for a quarterly P&L: two operators (one high-
margin, one mid-margin), fixed staff + infra
costs. Use placeholders like [Op A] and [X%].
You fill in the real numbers yourself, in your own document. The answer quality barely drops — try it in the simulator below.
Open the 🤖 AI Paste Simulator. Pick the "Quarterly P&L" sample, send it to the public chatbot (personal account) and watch what leaves the building. Then hit Sanitize and compare the answer you'd get. Finally switch to the approved internal tool and see why that path is fine even for the real thing.
5 · So when IS it fine to use AI?
Most of the time! This module should leave you using AI more confidently, not less:
✅ Approved internal / business-tier tools (ask IT which are approved for which data class): genuinely fine and encouraged — that's what they're contracted for.
✅ Public tools with non-sensitive content: "explain this Excel formula," "improve this generic paragraph," "brainstorm workshop names" — go for it.
✅ Public tools with sanitized prompts: structure and placeholders, no real names or numbers.
🚫 Public or personal-account tools with anything on the NEVER list — sanitize or switch tools.
Already pasted something you shouldn't have? Same rule as every module: report it fast — email infosec@tomhorngaming.com and tell your manager. Some vendors can delete conversations on request; the response team knows the routes. Reporting is always right and never punished (Module 8 has the full picture).
6 · Hands-on
Take one real prompt you'd genuinely use at work this week (or invent one with fake data). Write two versions: the raw paste and the sanitized rewrite with placeholders. Run the sanitized one through your approved tool — and notice the answer is just as useful. That rewrite reflex is the whole module.
7 · Quiz — 10 questions, pass at 8
Scenario-based, unlimited retries, every answer explained. No record of failed attempts is kept.