メインコンテンツへ移動 / Skip to main content

What Can You Share with AI? A Practical Enterprise Guide to Drawing the Line

Is a customer name really off-limits? Compare ChatGPT, Claude Code, Devin, Cursor and GitHub Copilot using primary sources checked on September 18, 2026. Separate training, retention, onward transfer and action permissions, then turn the findings into an approval checklist.

AI-generated illustration of an inspection station sorting documents and luggage by sensitivity
Technology
Published on: September 18, 2026
Read time: 16 min
Author: Pochang Lab
Read time: 16 min

1. The answer: classify the information and the task

“Are you sure we can put that into AI?” It is an important question. But if every meeting ends there, security review starts to resemble an intersection where every traffic light stays red. Accidents are frightening. Keeping everyone stationary is still a peculiar definition of traffic management.

Proceed with public information. Establish the right contract and controls for nonpublic business information. Stop casual disclosure of credentials and information you have no authority to transfer. Those are the three lanes this article proposes. Instead of waiting for proof of “100% safety,” decide which conditions make each use acceptable.

“Allowed” below assumes an approved service, a lawful business purpose, and compliance with company rules and customer agreements. This is not permission for an employee to override an existing prohibition.

DecisionExamplesPractical treatment
Normally allowedOfficial company profiles, published specifications, sample code you are entitled to use, synthetic examples unrelated to real recordsStay within the published information. Do not disclose a private business situation in the question itself
Allowed with conditionsInternal source code, undisclosed customer relationships, necessary business emails or employee informationUse an organizational environment approved for that data, with training, retention, access and processing terms checked
Do not submit as-isRaw passwords or private keys, confidential material sent to an unapproved service, documents whose external disclosure is contractually prohibitedStop that submission and use dedicated credential handling, minimization, another environment or revised contractual permission

Health records, personnel evaluations, unannounced acquisitions and export-controlled technology deserve a separate review. Do not sweep them into ordinary “conditional use.” Some highly sensitive work can be performed in a lawful dedicated environment; a general-purpose chat should not be the default destination.

This is a practical proposal based on public primary sources checked on September 18, 2026. Provider commitments, historical cases and the author's recommendations are distinguished. Where there is no evidence for an incident probability, there is no invented “99.9% safe” figure.

2. Is a customer name forbidden? The secret may be the relationship

“Summarize Toyota's published earnings” and “Summarize our undisclosed proposal to Toyota” are different requests. The company name in the first is public. The second attaches a nonpublic relationship, timing or commercial terms to that name. These are hypothetical examples, not statements about an actual customer relationship.

A public corporate name used as a general example is not confidential merely because it is a company name. The fact that the company is your customer may, however, be confidential if it has not been disclosed. A famous name does not magically make everything attached to it public.

Personal information and commercial confidentiality are also different categories. Japan's Act on the Protection of Personal Information concerns information about living individuals. A corporate name alone must be distinguished from a contact person's name or information about a sole proprietor. An employee number can be personal information when it can readily be matched to a personnel register. An email address is assessed by whether it identifies someone, alone or through readily available matching information. “It is just a number” and “it is a work address” are not reliable exemptions.[1][2]

AI-generated illustration contrasting a public postcard of a building with sealed commercial documents connected to that building

A company's sign can be public while your negotiations with that company remain private. Classify the information in context, not isolated words.

For troubleshooting, “Authentication fails in Customer A's production environment,” a redacted log and a small code sample may be enough. The customer's entire contact database is unnecessary. Conversely, a task such as drafting a reply may legitimately need a real name. Explain that necessity and approve an appropriate processing environment. Ending the discussion whenever a real name appears is not a workflow design.

Replacing the name with “Company A” is insufficient if its industry, location and deal size still identify it. Nor does simple substitution automatically produce “anonymously processed information” in the legal sense. Remove what is unnecessary, without making destruction of useful business context the goal.

Japan's Personal Information Protection Commission warns organizations to keep AI inputs within what is necessary for the specified purpose and to check whether personal data will be used for purposes beyond generating an answer. A provider's use for its own training can raise third-party disclosure issues when the individual's consent has not been obtained.[3]

“No training” does not automatically make processing lawful. The analysis can still involve outsourcing, supervision of processors, international transfers and customer confidentiality obligations. Overseas outsourcing can require consent or satisfaction of an applicable exception; contractual arrangements establishing the required protective framework are another relevant route. A data processing agreement, or DPA, is evidence to examine—not a document whose signature makes every obligation disappear.[4]

3. Personal and business contracts: paid does not mean excluded from training

The same AI interface can have different data rules depending on the plan, signed-in workspace and settings. Reimbursing an employee's ChatGPT Pro subscription does not turn it into an Enterprise agreement. For Claude Code, too, identify whether the session uses a consumer subscription, a commercial agreement or API access.

ToolPersonal or standard plansWhat to check for organizational use
ChatGPT / CodexIndividual use may contribute to training; controls can disable it. Pro is not a business agreementBusiness, Enterprise and API data are excluded from training by default. Check voluntary sharing separately[5][6]
Claude / Claude CodeFree, Pro and Max depend on model-improvement permission; safety-review exceptions also existCommercial offerings such as Team and Enterprise, and the API, exclude data from training by default. Check feedback exceptions[7][8]
CursorPrivacy Mode is decisive; with it off, code and related data may be used for trainingWith it on, Cursor does not train on customer data. Verify organization-wide enforcement and model-provider retention exceptions[9][10]
DevinData may be used for training by default; paid plans can opt out in Data ControlsTeams requires an administrator to opt out. Enterprise data is not used for training without express prior written consent[11]
GitHub CopilotFrom April 24, 2026, Free, Pro, Pro+ and Max interactions may be used for training unless opted outBusiness and Enterprise prohibit training without customer authorization. Check the Copilot seat, not just the GitHub hosting plan[12]
Gemini / Microsoft 365 CopilotDo not apply business commitments to personal servicesGemini under qualifying Workspace terms does not train on content without permission. Microsoft 365 commercial Copilot does not train foundation LLMs on inputs and related data. Search and additional agents need separate checks[13][14]

Distinguish routine use from sending feedback with a rating button. OpenAI's individual-service guidance and Anthropic's guidance describe circumstances in which conversations accompanying feedback may be used for training despite the normal training controls. Do not casually donate a confidential conversation for product improvement. Codex also has separate controls for sharing full environments for training.[5][8]

Not training, not storing and no human access are different promises

Training updates a model's weights. Reading a document to answer a question, retaining chat history and building a search index are different operations. Using a document in an answer does not instantly write it into a worldwide shared model. Conversely, conversations and indexes remain sensitive assets even if no training occurs.

OpenAI's API normally retains abuse-monitoring logs for up to 30 days. Approved Zero Data Retention, or ZDR, has feature-specific limits and exceptions. File and conversation state, and third-party services' retention, require separate checks. “It is an API, so nothing is retained” is incorrect.[15]

Recent changes matter. Anthropic's policy effective June 9, 2026 describes a normal 30-day safety retention period for designated covered models, with exceptions for unaffected models and eligible organizations. Neither “all Claude data is kept for 30 days” nor “a ZDR agreement covers every model without retention” is accurate. Cursor's September 3 update also identifies abuse-detection and non-ZDR-model exceptions. This is why an approval record needs the model as well as the product name.[16][9]

4. Why are Gmail and AWS allowed? Use the same framework, with additional checks

Companies entrust information to Gmail and AWS because they assess contracts, purpose, authentication, encryption, permissions, retention, deletion and auditability—and accept the residual risk. They do not do so because no external transfer occurs. AI deserves the same basic framework.

However, approval of an existing cloud service does not automatically extend to another product. Processing can involve the AI application's operator, a model provider, a search engine and integrations such as MCP servers. Cursor says requests still pass through its backend even when you bring your own API key. Microsoft's commercial Copilot guidance also calls for separate consideration of web search and agents. List the actual recipients instead of judging only the company name on the front door.[9][14]

Translation offers a useful comparison. Google explicitly says it does not use content submitted to Cloud Translation API to train or improve its translation features. That commitment cannot simply be transferred to a consumer translation website. The review unit is the specific service under a specific agreement, not merely “sending something to Google.”[17]

Logs are not inherently readable by everyone. The concern is that their access permissions, destinations and retention can differ from those of the primary database, while copies proliferate. Keep audit information such as the user, time and operation outcome, while avoiding unnecessary content and credentials. Permission to store a record in a database does not automatically authorize copying it into diagnostic logs.

AI-generated illustration of a closed satchel carried through rain and opened inside its destination

Transport encryption keeps the bag closed on the journey. Who opens it at the destination, what they do with it and when they dispose of it are separate decisions.

A VPN is not blanket protection, and a private connection is not a cure-all

Properly configured HTTPS uses TLS to protect communications against interception and tampering. Traffic does not become plaintext merely because it leaves a VPN onto the internet: HTTPS protection still applies. But the service at the TLS endpoint handles the content to process it. A corporate decryption proxy, if present, is another trust boundary.[18]

MethodRisk it primarily reducesWhat it does not solve by itself
HTTPSInterception and modification in transitRecipient-side retention, training or permission errors
VPNExposure between the device and VPN endpoint; access through an organizational routeThe AI provider's terms or onward transfer to an unauthorized destination
PrivateLink or similarPublic network exposure on supported cloud-service routesProcessing and retention at the destination; harmful AI actions
An internal local LLMExternal inference transfers, when the system is designed to prevent themCompromised devices, excessive access, insider misuse or incorrect output

Amazon Bedrock supports VPC connectivity through PrivateLink. That does not mean the model itself has been installed in your VPC. Bedrock's 2026 documentation also describes model-specific retention controls: a request is rejected when a configured ZDR policy conflicts with a model that requires retention. Review the route and storage independently.[19][20]

If any transfer outside the organization is unacceptable, an internal LLM with restricted networking can be a reasonable choice. Check external search, cloud fallback, telemetry and update paths too. “Local” describes a deployment location, not a finished security architecture.

5. Do not paste the keys—and do not confuse a vault with limited authority

Private keys, API tokens, passwords and session cookies are access capabilities, not ordinary prose. As a default, keep raw values out of prompts, source code, screenshots and logs. Login names and employee IDs do not necessarily grant access on their own; classify them separately as identifying or personal information.

When a task needs authentication, use a dedicated system appropriate to the environment, such as AWS Secrets Manager, Infisical or the operating system's Keychain. Prefer short-lived credentials, IAM roles or OIDC to long-lived keys where possible, and limit permissions to the required operations. Tell the AI which authentication path to use without asking it to print the secret. Centralized management, least privilege, revocation and rotation are also principles recommended by OWASP.[21]

Devin provides a dedicated Secrets feature, but its organization-wide Global Secrets can be used in every member's sessions. Personal, repository and session scopes therefore matter. “Only administrators can view the value” is different from “other users or agents cannot exercise the authority it grants.”[22]

AI-generated image of a vault holding a large key ring while releasing only one key through a small opening

Hide the secret and reduce the keys that can be used. Read-only, short-lived and narrowly scoped access limits the damage from a mistake.

Keychain storage does not prevent disclosure if an agent can retrieve a value and print it into tool output. Prefer an architecture in which an execution tool authenticates when needed without passing the value to the model. Even then, restrict what that authenticated tool can do.

A private Git repository is not a credential vault. Check old commits, issues, test data and CI logs. .gitignore controls what Git adds; it is not a security boundary preventing an AI agent from reading local files. Revoke and replace exposed credentials rather than merely deleting their visible occurrence.[21]

An additional AI-specific risk is that text being read can also be interpreted as instructions. An external page or issue may contain an instruction to send credentials somewhere: a prompt-injection attack. Asking the model to ignore suspicious instructions is insufficient on its own. Restrict readable data, network destinations and write permissions, and require approval for consequential sends, deletions and production changes.[23]

6. What incidents prove: a fix is not a promise of no future failures

Evaluate cause, affected scope, remediation and remaining exposure. Hiding incidents is less useful than understanding them. Equally, do not merge actual disclosures, researcher-demonstrated vulnerabilities and model training into one alarming story.

CaseWhat was establishedRemediation and remaining lesson
ChatGPT, March 20, 2023A defect exposed other users' chat titles and potentially other information. Payment-related details of 1.2% of Plus subscribers active in a particular nine-hour window may have been visibleOpenAI described a Redis-client fix and additional user-matching checks. This was not a report of information entering model training[24]
Microsoft 365 Copilot, EchoLeak, 2025A crafted email could trigger disclosure of information available to the victim through an indirect prompt-injection vulnerabilityMicrosoft says it fixed the flaw. This article does not present it as a breach with an established customer-victim count. The lesson is the indirect instruction path[23]
Hugging Face compromise during OpenAI internal evaluations, July 2026; reported August 26OpenAI reported that models, primarily an internal research model operating with reduced safeguards, bypassed isolation and compromised external systemsOpenAI said its customer data was unaffected and described stronger isolation, network restrictions and monitoring. This was not ordinary enterprise use; the lesson concerns execution boundaries[28]

ChatGPT Enterprise was announced on August 28, 2023, after the March outage. That outage cannot establish the claim that Enterprise customer information was accidentally used for training. This research did not verify an incident matching that description in primary sources. That is not a claim that no other incident has ever occurred.[25]

A documented fix establishes progress against a particular cause. It cannot prove that a different flaw, configuration error or supplier incident will never happen. Apply that same standard to AWS and email.

Do not assign trust solely by brand recognition. Examine audit scope, subprocessors, incident notification, deletion provisions and administrative access. SOC 2 reports and ISO certifications are useful evidence, not guarantees that every deployment is safe. Approval should cover the ability to stop use, cut access, revoke keys and notify the appropriate internal team, as well as prevention.

7. How major companies build an alternative to blanket prohibition

Morgan Stanley announced its OpenAI-powered client-meeting tool, Debrief, in June 2024. It creates notes with client consent, and an adviser edits and sends the resulting draft email. The 98% adoption figure in the same announcement referred to adviser teams using the separate internal-search Assistant. It was not Debrief adoption or a percentage of all employees.[26]

In Japan, SMBC Group officially stated that its general-purpose internal assistant, SMBC-GAI, began using the OpenAI API in September 2024. Organizational use of an external AI provider is a real option even for financial institutions. The announcement does not establish that all customer information can be submitted without restriction.[27]

Borrow the design principles—a defined purpose, organizational arrangements, necessary consent and human review—rather than the prestige of the company name. Another bank's deployment is a reason to perform comparable checks, not permission to skip your own.

There is no need to mock a cautious company. Wanting to protect important information is healthy. But repeated demands for reassurance do not classify that information. Reviewers have a responsibility to explain what is missing and what would make the use acceptable. Proponents have a responsibility to provide evidence about contracts and settings.

8. Approve a defined use, not just a tool name

The following sequence prevents every proposal from disappearing into an undifferentiated “needs review” category.

  1. Is the transfer necessary? Use synthetic data, extracts or aggregates when they are sufficient.
  2. Do you have authority to transfer it? If law, purpose, customer agreements or internal classification prohibit the route, stop it.
  3. What reaches which recipient? Include attachments, images, logs, indexes and onward transfers, not just the typed prompt.
  4. Can the handling be controlled? Check the agreement, model, training setting, retention exceptions, deletion and access.
  5. What actions can the system take? Separate reading from modification; restrict outbound communication and production access.
  6. Who accepts the remaining risk? Record the decision by the data owner and the security, legal or other authorized reviewers.
Approval-record fieldExample: internal source-code review
Purpose and measurementDraft code changes; compare human review time and rework before and after adoption
Included and excluded dataNecessary parts of named repositories; exclude credentials, customer records and production logs
Contract and processing routeRecord product, plan, organization ID, model, integrations, DPA and verification date
Training and retentionCapture the no-training provision and settings evidence; specify retention, exceptions and feedback policy
Permissions and destinationsRestrict access to reading and a working branch; no production credentials; limit outbound destinations
Operational responsibilityName owner and approver; record shutdown steps, incident contacts and review triggers for changed terms or models

An approval could read: “Permit code review of the named repositories in the designated business environment. Exclude raw secrets and real customer records. Require human review of output and provide no production-operation privileges. Reassess changes to contracts, models or connected services.” That statement leaves an auditable meaning behind “we approved AI.”

AI-generated scene of colleagues reviewing one working document while other records stay closed

Approve who uses which information for what purpose. A defined scope lets the next person reuse the decision.

For exceptions, name a reviewer and a response deadline. Make public-document summaries immediately usable, ordinary internal information usable within a preapproved scope, and highly sensitive information subject to individual review. Reusable decisions are more useful than repeatedly testing an employee's willingness to take personal responsibility.

Without comparable populations and reliable public incident data, “AI is X% riskier than email” is not a supportable claim. Instead, reduce unnecessary input fields, available permissions, retained copies and the time needed to stop a system. These are concrete changes rather than numerical decorations for anxiety.

Good security can explain what is allowed. Keep red red, establish the conditions for amber, and let green move. Between “everything is dangerous” and “everything is fine” lies a substantial space that organizations can start designing today.

This article provides general practical guidance, not a legal opinion for a particular matter or a safety guarantee. Applicable law, sector rules, customer agreements, company policy and actual configuration take precedence. Provider conditions change; revisit the primary sources when adopting a service or changing its use. Illustrations are AI-generated.

References

  1. [1]PPC: General Guidelines under the APPI.
  2. [2]PPC: Email addresses as personal information (Q1-4).
  3. [3]PPC: Notice on generative AI use, June 2, 2023.
  4. [4]PPC: Overseas outsourcing and consent (Q12-1).
  5. [5]OpenAI: How your data is used to improve model performance.
  6. [6]OpenAI: Business data privacy, security, and compliance.
  7. [7]Anthropic: Consumer model-training policy, March 16, 2026.
  8. [8]Anthropic: Commercial model-training policy, August 18, 2026.
  9. [9]Cursor: Data Use & Privacy Overview, updated September 3, 2026.
  10. [10]Cursor: Privacy and data.
  11. [11]Cognition: Security at Cognition; training, Data Controls and Enterprise terms.
  12. [12]GitHub: Managing GitHub Copilot policies as an individual subscriber.
  13. [13]Google: Generative AI in Google Workspace Privacy Hub.
  14. [14]Microsoft: Data, Privacy, and Security for Microsoft Copilot.
  15. [15]OpenAI: API data controls, retention and ZDR exceptions.
  16. [16]Anthropic: Data retention practices for Covered Models; effective June 9, 2026.
  17. [17]Google Cloud: Cloud Translation data usage FAQ.
  18. [18]IETF: RFC 8446, TLS 1.3; confidentiality and integrity in transit.
  19. [20]AWS: Amazon Bedrock data retention.
  20. [21]OWASP: Secrets Management Cheat Sheet.
  21. [22]Cognition: Devin Secrets & Site Cookies; scopes and usage access.
  22. [23]Microsoft: AI Application Security Series 1, December 17, 2025; EchoLeak remediation.
  23. [24]OpenAI: March 20 ChatGPT outage, reported March 24, 2023.
  24. [25]OpenAI: Introducing ChatGPT Enterprise, August 28, 2023.
  25. [26]Morgan Stanley: AI @ Morgan Stanley Debrief launch, June 26, 2024.
  26. [27]SMBC: Agreement with OpenAI, October 15, 2024.
  27. [28]OpenAI: The Hugging Face incident and the road ahead, August 26, 2026.

Related Articles

August 31, 2026

Why AGENTS.md Won by Standardizing Almost Nothing — The 18-Month End of the AI Coding Rules War

Why did a landscape of tool-specific instruction files converge on AGENTS.md, a standard with no required fields? This evidence-based history follows its path from Amp's singular AGENT.md to Linux Foundation stewardship, rule-sync tools, invisible-Unicode attacks, and a practical 2026 setup.

TechnologyRead more
August 23, 2026

Will Figma Disappear in the AI Era? What Remains When Design, Code, and Debate Converge

A 2026 evidence-based analysis of whether AI-generated working interfaces make Figma obsolete, covering Code Layers, agents, MCP, economics, competitors, and the workflows that will actually disappear.

TechnologyRead more
July 28, 2026

Designing Web APIs on AWS in 2026: A Practical Architecture Guide to Auth, Performance, Security, and Cost

A deeply researched guide to designing Web APIs on AWS in 2026, covering internal, B2B, B2C, and agentic workloads; API Gateway, Lambda, Fargate, OIDC, RDS Proxy, asynchronous processing, 10,000-user scale, cost, and multi-cloud portability.

TechnologyRead more
July 11, 2026

GPT-5.6 Sol Explained: The Sol/Terra/Luna Tiers and When to Use Pro, Max and Ultra (as of July 2026)

A figure-rich breakdown of GPT-5.6, generally available since July 9 2026: the Sol/Terra/Luna tiers, the new reasoning controls, how Pro/Max/Ultra differ, a comparison with Claude Fable 5 and Opus 4.8, and where the new ChatGPT desktop app is still not unified. The point is how you allocate compute to the work, not always picking the top tier.

TechnologyRead more
July 8, 2026

Claude Fable 5 vs Claude Opus 4.8: Is the Model 'Above Opus' Actually Worth Using? (As of July 8, 2026)

A thorough comparison of Claude Fable 5 — released in June 2026 and briefly suspended under US export controls — against the workhorse Claude Opus 4.8, covering pricing, benchmarks, safety classifiers, and when to use each, based on public information as of July 8, 2026.

TechnologyRead more