Table of Contents
1. The answer: classify the information and the task
“Are you sure we can put that into AI?” It is an important question. But if every meeting ends there, security review starts to resemble an intersection where every traffic light stays red. Accidents are frightening. Keeping everyone stationary is still a peculiar definition of traffic management.
Proceed with public information. Establish the right contract and controls for nonpublic business information. Stop casual disclosure of credentials and information you have no authority to transfer. Those are the three lanes this article proposes. Instead of waiting for proof of “100% safety,” decide which conditions make each use acceptable.
“Allowed” below assumes an approved service, a lawful business purpose, and compliance with company rules and customer agreements. This is not permission for an employee to override an existing prohibition.
| Decision | Examples | Practical treatment |
|---|---|---|
| Normally allowed | Official company profiles, published specifications, sample code you are entitled to use, synthetic examples unrelated to real records | Stay within the published information. Do not disclose a private business situation in the question itself |
| Allowed with conditions | Internal source code, undisclosed customer relationships, necessary business emails or employee information | Use an organizational environment approved for that data, with training, retention, access and processing terms checked |
| Do not submit as-is | Raw passwords or private keys, confidential material sent to an unapproved service, documents whose external disclosure is contractually prohibited | Stop that submission and use dedicated credential handling, minimization, another environment or revised contractual permission |
Health records, personnel evaluations, unannounced acquisitions and export-controlled technology deserve a separate review. Do not sweep them into ordinary “conditional use.” Some highly sensitive work can be performed in a lawful dedicated environment; a general-purpose chat should not be the default destination.
This is a practical proposal based on public primary sources checked on September 18, 2026. Provider commitments, historical cases and the author's recommendations are distinguished. Where there is no evidence for an incident probability, there is no invented “99.9% safe” figure.
2. Is a customer name forbidden? The secret may be the relationship
“Summarize Toyota's published earnings” and “Summarize our undisclosed proposal to Toyota” are different requests. The company name in the first is public. The second attaches a nonpublic relationship, timing or commercial terms to that name. These are hypothetical examples, not statements about an actual customer relationship.
A public corporate name used as a general example is not confidential merely because it is a company name. The fact that the company is your customer may, however, be confidential if it has not been disclosed. A famous name does not magically make everything attached to it public.
Personal information and commercial confidentiality are also different categories. Japan's Act on the Protection of Personal Information concerns information about living individuals. A corporate name alone must be distinguished from a contact person's name or information about a sole proprietor. An employee number can be personal information when it can readily be matched to a personnel register. An email address is assessed by whether it identifies someone, alone or through readily available matching information. “It is just a number” and “it is a work address” are not reliable exemptions.[1][2]
A company's sign can be public while your negotiations with that company remain private. Classify the information in context, not isolated words.
For troubleshooting, “Authentication fails in Customer A's production environment,” a redacted log and a small code sample may be enough. The customer's entire contact database is unnecessary. Conversely, a task such as drafting a reply may legitimately need a real name. Explain that necessity and approve an appropriate processing environment. Ending the discussion whenever a real name appears is not a workflow design.
Replacing the name with “Company A” is insufficient if its industry, location and deal size still identify it. Nor does simple substitution automatically produce “anonymously processed information” in the legal sense. Remove what is unnecessary, without making destruction of useful business context the goal.
Japan's Personal Information Protection Commission warns organizations to keep AI inputs within what is necessary for the specified purpose and to check whether personal data will be used for purposes beyond generating an answer. A provider's use for its own training can raise third-party disclosure issues when the individual's consent has not been obtained.[3]
“No training” does not automatically make processing lawful. The analysis can still involve outsourcing, supervision of processors, international transfers and customer confidentiality obligations. Overseas outsourcing can require consent or satisfaction of an applicable exception; contractual arrangements establishing the required protective framework are another relevant route. A data processing agreement, or DPA, is evidence to examine—not a document whose signature makes every obligation disappear.[4]
3. Personal and business contracts: paid does not mean excluded from training
The same AI interface can have different data rules depending on the plan, signed-in workspace and settings. Reimbursing an employee's ChatGPT Pro subscription does not turn it into an Enterprise agreement. For Claude Code, too, identify whether the session uses a consumer subscription, a commercial agreement or API access.
| Tool | Personal or standard plans | What to check for organizational use |
|---|---|---|
| ChatGPT / Codex | Individual use may contribute to training; controls can disable it. Pro is not a business agreement | Business, Enterprise and API data are excluded from training by default. Check voluntary sharing separately[5][6] |
| Claude / Claude Code | Free, Pro and Max depend on model-improvement permission; safety-review exceptions also exist | Commercial offerings such as Team and Enterprise, and the API, exclude data from training by default. Check feedback exceptions[7][8] |
| Cursor | Privacy Mode is decisive; with it off, code and related data may be used for training | With it on, Cursor does not train on customer data. Verify organization-wide enforcement and model-provider retention exceptions[9][10] |
| Devin | Data may be used for training by default; paid plans can opt out in Data Controls | Teams requires an administrator to opt out. Enterprise data is not used for training without express prior written consent[11] |
| GitHub Copilot | From April 24, 2026, Free, Pro, Pro+ and Max interactions may be used for training unless opted out | Business and Enterprise prohibit training without customer authorization. Check the Copilot seat, not just the GitHub hosting plan[12] |
| Gemini / Microsoft 365 Copilot | Do not apply business commitments to personal services | Gemini under qualifying Workspace terms does not train on content without permission. Microsoft 365 commercial Copilot does not train foundation LLMs on inputs and related data. Search and additional agents need separate checks[13][14] |
Distinguish routine use from sending feedback with a rating button. OpenAI's individual-service guidance and Anthropic's guidance describe circumstances in which conversations accompanying feedback may be used for training despite the normal training controls. Do not casually donate a confidential conversation for product improvement. Codex also has separate controls for sharing full environments for training.[5][8]
Not training, not storing and no human access are different promises
Training updates a model's weights. Reading a document to answer a question, retaining chat history and building a search index are different operations. Using a document in an answer does not instantly write it into a worldwide shared model. Conversely, conversations and indexes remain sensitive assets even if no training occurs.
OpenAI's API normally retains abuse-monitoring logs for up to 30 days. Approved Zero Data Retention, or ZDR, has feature-specific limits and exceptions. File and conversation state, and third-party services' retention, require separate checks. “It is an API, so nothing is retained” is incorrect.[15]
Recent changes matter. Anthropic's policy effective June 9, 2026 describes a normal 30-day safety retention period for designated covered models, with exceptions for unaffected models and eligible organizations. Neither “all Claude data is kept for 30 days” nor “a ZDR agreement covers every model without retention” is accurate. Cursor's September 3 update also identifies abuse-detection and non-ZDR-model exceptions. This is why an approval record needs the model as well as the product name.[16][9]
4. Why are Gmail and AWS allowed? Use the same framework, with additional checks
Companies entrust information to Gmail and AWS because they assess contracts, purpose, authentication, encryption, permissions, retention, deletion and auditability—and accept the residual risk. They do not do so because no external transfer occurs. AI deserves the same basic framework.
However, approval of an existing cloud service does not automatically extend to another product. Processing can involve the AI application's operator, a model provider, a search engine and integrations such as MCP servers. Cursor says requests still pass through its backend even when you bring your own API key. Microsoft's commercial Copilot guidance also calls for separate consideration of web search and agents. List the actual recipients instead of judging only the company name on the front door.[9][14]
Translation offers a useful comparison. Google explicitly says it does not use content submitted to Cloud Translation API to train or improve its translation features. That commitment cannot simply be transferred to a consumer translation website. The review unit is the specific service under a specific agreement, not merely “sending something to Google.”[17]
Logs are not inherently readable by everyone. The concern is that their access permissions, destinations and retention can differ from those of the primary database, while copies proliferate. Keep audit information such as the user, time and operation outcome, while avoiding unnecessary content and credentials. Permission to store a record in a database does not automatically authorize copying it into diagnostic logs.
Transport encryption keeps the bag closed on the journey. Who opens it at the destination, what they do with it and when they dispose of it are separate decisions.
A VPN is not blanket protection, and a private connection is not a cure-all
Properly configured HTTPS uses TLS to protect communications against interception and tampering. Traffic does not become plaintext merely because it leaves a VPN onto the internet: HTTPS protection still applies. But the service at the TLS endpoint handles the content to process it. A corporate decryption proxy, if present, is another trust boundary.[18]
| Method | Risk it primarily reduces | What it does not solve by itself |
|---|---|---|
| HTTPS | Interception and modification in transit | Recipient-side retention, training or permission errors |
| VPN | Exposure between the device and VPN endpoint; access through an organizational route | The AI provider's terms or onward transfer to an unauthorized destination |
| PrivateLink or similar | Public network exposure on supported cloud-service routes | Processing and retention at the destination; harmful AI actions |
| An internal local LLM | External inference transfers, when the system is designed to prevent them | Compromised devices, excessive access, insider misuse or incorrect output |
Amazon Bedrock supports VPC connectivity through PrivateLink. That does not mean the model itself has been installed in your VPC. Bedrock's 2026 documentation also describes model-specific retention controls: a request is rejected when a configured ZDR policy conflicts with a model that requires retention. Review the route and storage independently.[19][20]
If any transfer outside the organization is unacceptable, an internal LLM with restricted networking can be a reasonable choice. Check external search, cloud fallback, telemetry and update paths too. “Local” describes a deployment location, not a finished security architecture.
5. Do not paste the keys—and do not confuse a vault with limited authority
Private keys, API tokens, passwords and session cookies are access capabilities, not ordinary prose. As a default, keep raw values out of prompts, source code, screenshots and logs. Login names and employee IDs do not necessarily grant access on their own; classify them separately as identifying or personal information.
When a task needs authentication, use a dedicated system appropriate to the environment, such as AWS Secrets Manager, Infisical or the operating system's Keychain. Prefer short-lived credentials, IAM roles or OIDC to long-lived keys where possible, and limit permissions to the required operations. Tell the AI which authentication path to use without asking it to print the secret. Centralized management, least privilege, revocation and rotation are also principles recommended by OWASP.[21]
Devin provides a dedicated Secrets feature, but its organization-wide Global Secrets can be used in every member's sessions. Personal, repository and session scopes therefore matter. “Only administrators can view the value” is different from “other users or agents cannot exercise the authority it grants.”[22]
Hide the secret and reduce the keys that can be used. Read-only, short-lived and narrowly scoped access limits the damage from a mistake.
Keychain storage does not prevent disclosure if an agent can retrieve a value and print it into tool output. Prefer an architecture in which an execution tool authenticates when needed without passing the value to the model. Even then, restrict what that authenticated tool can do.
A private Git repository is not a credential vault. Check old commits, issues, test data and CI logs. .gitignore controls what Git adds; it is not a security boundary preventing an AI agent from reading local files. Revoke and replace exposed credentials rather than merely deleting their visible occurrence.[21]
An additional AI-specific risk is that text being read can also be interpreted as instructions. An external page or issue may contain an instruction to send credentials somewhere: a prompt-injection attack. Asking the model to ignore suspicious instructions is insufficient on its own. Restrict readable data, network destinations and write permissions, and require approval for consequential sends, deletions and production changes.[23]
6. What incidents prove: a fix is not a promise of no future failures
Evaluate cause, affected scope, remediation and remaining exposure. Hiding incidents is less useful than understanding them. Equally, do not merge actual disclosures, researcher-demonstrated vulnerabilities and model training into one alarming story.
| Case | What was established | Remediation and remaining lesson |
|---|---|---|
| ChatGPT, March 20, 2023 | A defect exposed other users' chat titles and potentially other information. Payment-related details of 1.2% of Plus subscribers active in a particular nine-hour window may have been visible | OpenAI described a Redis-client fix and additional user-matching checks. This was not a report of information entering model training[24] |
| Microsoft 365 Copilot, EchoLeak, 2025 | A crafted email could trigger disclosure of information available to the victim through an indirect prompt-injection vulnerability | Microsoft says it fixed the flaw. This article does not present it as a breach with an established customer-victim count. The lesson is the indirect instruction path[23] |
| Hugging Face compromise during OpenAI internal evaluations, July 2026; reported August 26 | OpenAI reported that models, primarily an internal research model operating with reduced safeguards, bypassed isolation and compromised external systems | OpenAI said its customer data was unaffected and described stronger isolation, network restrictions and monitoring. This was not ordinary enterprise use; the lesson concerns execution boundaries[28] |
ChatGPT Enterprise was announced on August 28, 2023, after the March outage. That outage cannot establish the claim that Enterprise customer information was accidentally used for training. This research did not verify an incident matching that description in primary sources. That is not a claim that no other incident has ever occurred.[25]
A documented fix establishes progress against a particular cause. It cannot prove that a different flaw, configuration error or supplier incident will never happen. Apply that same standard to AWS and email.
Do not assign trust solely by brand recognition. Examine audit scope, subprocessors, incident notification, deletion provisions and administrative access. SOC 2 reports and ISO certifications are useful evidence, not guarantees that every deployment is safe. Approval should cover the ability to stop use, cut access, revoke keys and notify the appropriate internal team, as well as prevention.
7. How major companies build an alternative to blanket prohibition
Morgan Stanley announced its OpenAI-powered client-meeting tool, Debrief, in June 2024. It creates notes with client consent, and an adviser edits and sends the resulting draft email. The 98% adoption figure in the same announcement referred to adviser teams using the separate internal-search Assistant. It was not Debrief adoption or a percentage of all employees.[26]
In Japan, SMBC Group officially stated that its general-purpose internal assistant, SMBC-GAI, began using the OpenAI API in September 2024. Organizational use of an external AI provider is a real option even for financial institutions. The announcement does not establish that all customer information can be submitted without restriction.[27]
Borrow the design principles—a defined purpose, organizational arrangements, necessary consent and human review—rather than the prestige of the company name. Another bank's deployment is a reason to perform comparable checks, not permission to skip your own.
There is no need to mock a cautious company. Wanting to protect important information is healthy. But repeated demands for reassurance do not classify that information. Reviewers have a responsibility to explain what is missing and what would make the use acceptable. Proponents have a responsibility to provide evidence about contracts and settings.
8. Approve a defined use, not just a tool name
The following sequence prevents every proposal from disappearing into an undifferentiated “needs review” category.
- Is the transfer necessary? Use synthetic data, extracts or aggregates when they are sufficient.
- Do you have authority to transfer it? If law, purpose, customer agreements or internal classification prohibit the route, stop it.
- What reaches which recipient? Include attachments, images, logs, indexes and onward transfers, not just the typed prompt.
- Can the handling be controlled? Check the agreement, model, training setting, retention exceptions, deletion and access.
- What actions can the system take? Separate reading from modification; restrict outbound communication and production access.
- Who accepts the remaining risk? Record the decision by the data owner and the security, legal or other authorized reviewers.
| Approval-record field | Example: internal source-code review |
|---|---|
| Purpose and measurement | Draft code changes; compare human review time and rework before and after adoption |
| Included and excluded data | Necessary parts of named repositories; exclude credentials, customer records and production logs |
| Contract and processing route | Record product, plan, organization ID, model, integrations, DPA and verification date |
| Training and retention | Capture the no-training provision and settings evidence; specify retention, exceptions and feedback policy |
| Permissions and destinations | Restrict access to reading and a working branch; no production credentials; limit outbound destinations |
| Operational responsibility | Name owner and approver; record shutdown steps, incident contacts and review triggers for changed terms or models |
An approval could read: “Permit code review of the named repositories in the designated business environment. Exclude raw secrets and real customer records. Require human review of output and provide no production-operation privileges. Reassess changes to contracts, models or connected services.” That statement leaves an auditable meaning behind “we approved AI.”
Approve who uses which information for what purpose. A defined scope lets the next person reuse the decision.
For exceptions, name a reviewer and a response deadline. Make public-document summaries immediately usable, ordinary internal information usable within a preapproved scope, and highly sensitive information subject to individual review. Reusable decisions are more useful than repeatedly testing an employee's willingness to take personal responsibility.
Without comparable populations and reliable public incident data, “AI is X% riskier than email” is not a supportable claim. Instead, reduce unnecessary input fields, available permissions, retained copies and the time needed to stop a system. These are concrete changes rather than numerical decorations for anxiety.
Good security can explain what is allowed. Keep red red, establish the conditions for amber, and let green move. Between “everything is dangerous” and “everything is fine” lies a substantial space that organizations can start designing today.
This article provides general practical guidance, not a legal opinion for a particular matter or a safety guarantee. Applicable law, sector rules, customer agreements, company policy and actual configuration take precedence. Provider conditions change; revisit the primary sources when adopting a service or changing its use. Illustrations are AI-generated.
References
- [1]PPC: General Guidelines under the APPI. ↩
- [2]PPC: Email addresses as personal information (Q1-4). ↩
- [3]PPC: Notice on generative AI use, June 2, 2023. ↩
- [4]PPC: Overseas outsourcing and consent (Q12-1). ↩
- [5]OpenAI: How your data is used to improve model performance. ↩
- [6]OpenAI: Business data privacy, security, and compliance. ↩
- [7]Anthropic: Consumer model-training policy, March 16, 2026. ↩
- [8]Anthropic: Commercial model-training policy, August 18, 2026. ↩
- [9]Cursor: Data Use & Privacy Overview, updated September 3, 2026. ↩
- [10]Cursor: Privacy and data. ↩
- [11]Cognition: Security at Cognition; training, Data Controls and Enterprise terms. ↩
- [12]GitHub: Managing GitHub Copilot policies as an individual subscriber. ↩
- [13]Google: Generative AI in Google Workspace Privacy Hub. ↩
- [14]Microsoft: Data, Privacy, and Security for Microsoft Copilot. ↩
- [15]OpenAI: API data controls, retention and ZDR exceptions. ↩
- [16]Anthropic: Data retention practices for Covered Models; effective June 9, 2026. ↩
- [17]Google Cloud: Cloud Translation data usage FAQ. ↩
- [18]IETF: RFC 8446, TLS 1.3; confidentiality and integrity in transit. ↩
- [19]AWS: Protect your data using Amazon VPC and AWS PrivateLink. ↩
- [20]AWS: Amazon Bedrock data retention. ↩
- [21]OWASP: Secrets Management Cheat Sheet. ↩
- [22]Cognition: Devin Secrets & Site Cookies; scopes and usage access. ↩
- [23]Microsoft: AI Application Security Series 1, December 17, 2025; EchoLeak remediation. ↩
- [24]OpenAI: March 20 ChatGPT outage, reported March 24, 2023. ↩
- [25]OpenAI: Introducing ChatGPT Enterprise, August 28, 2023. ↩
- [26]Morgan Stanley: AI @ Morgan Stanley Debrief launch, June 26, 2024. ↩
- [27]SMBC: Agreement with OpenAI, October 15, 2024. ↩
- [28]OpenAI: The Hugging Face incident and the road ahead, August 26, 2026. ↩

NEW NOVEL 2026/08/01
Clouded Glass
Polishing is not about force.
Volume two of The World Became Slightly Farther Away.Five stories that can also be read as a starting point.
View on Amazon
Jijoden.com
Your life is worth writing.
There is a truer self you can tell only to AI.Gather fragments of memory into a single story.
Take a LookRelated Articles
Why AGENTS.md Won by Standardizing Almost Nothing — The 18-Month End of the AI Coding Rules War
Why did a landscape of tool-specific instruction files converge on AGENTS.md, a standard with no required fields? This evidence-based history follows its path from Amp's singular AGENT.md to Linux Foundation stewardship, rule-sync tools, invisible-Unicode attacks, and a practical 2026 setup.
Will Figma Disappear in the AI Era? What Remains When Design, Code, and Debate Converge
A 2026 evidence-based analysis of whether AI-generated working interfaces make Figma obsolete, covering Code Layers, agents, MCP, economics, competitors, and the workflows that will actually disappear.
Designing Web APIs on AWS in 2026: A Practical Architecture Guide to Auth, Performance, Security, and Cost
A deeply researched guide to designing Web APIs on AWS in 2026, covering internal, B2B, B2C, and agentic workloads; API Gateway, Lambda, Fargate, OIDC, RDS Proxy, asynchronous processing, 10,000-user scale, cost, and multi-cloud portability.
GPT-5.6 Sol Explained: The Sol/Terra/Luna Tiers and When to Use Pro, Max and Ultra (as of July 2026)
A figure-rich breakdown of GPT-5.6, generally available since July 9 2026: the Sol/Terra/Luna tiers, the new reasoning controls, how Pro/Max/Ultra differ, a comparison with Claude Fable 5 and Opus 4.8, and where the new ChatGPT desktop app is still not unified. The point is how you allocate compute to the work, not always picking the top tier.
Claude Fable 5 vs Claude Opus 4.8: Is the Model 'Above Opus' Actually Worth Using? (As of July 8, 2026)
A thorough comparison of Claude Fable 5 — released in June 2026 and briefly suspended under US export controls — against the workhorse Claude Opus 4.8, covering pricing, benchmarks, safety classifiers, and when to use each, based on public information as of July 8, 2026.
