Stepping Up to the Plate: Public LLMs and Risks to Confidentiality

August 25, 2026

In our previous post in this series, we discussed cautionary tales that stem from the misuse of AI in case filings, what hallucinations really are, and why the use of AI differs for general purpose AI models like ChatGPT vs. purpose-built models for eDiscovery.

As AI model chatbots – like ChatGPT, Claude and Google Gemini – have become more ubiquitous, people have been more inclined to use these models for practically anything. That prolific use has extended to using the models in litigation.

We discussed in our previous post that many pro se parties are relying on public large language models (LLMs) like ChatGPT to draft legal filings, which is why a majority of AI hallucinations cases have been caused by pro se parties.

But drafting legal filings isn’t the only use case where pro se parties and even lawyers are using public LLMs in litigation. In some cases, they are using them to analyse documents produced in discovery by opposing counsel. In this post, we’ll discuss why the use of public LLMs on opponents’ documents is not a good idea, how the risk for public LLMs differs from the use of other technology platforms, and we’ll weigh the risk of inadvertent disclosure against the potential access to justice benefits of using lower cost AI solutions.

Why is the Use of Public LLMs an Issue?

In our previous post, we discussed that public AI models use broad training data and probabilistic language generation to answer virtually any question, which makes them more likely to hallucinate than eDiscovery AI solutions that are purpose-built and designed to work within a defined corpus of evidence.

The use of broad training data leads to a second concern regarding the use of public LLMs in litigation, their tendency (by default, at least) to use the chat histories of their users to continue to train their models. Uploading documents for analysis by public AI models, a common use case for many can lead to the data within those documents to be used for training the model. That data, or portions of it could then be included in an answer to another user in response to their prompt. If that data contains an organisation’s sensitive or confidential data, the result could be disastrous.

Additionally, the exposure of the data might not be limited to users of the model. A recent article in Futurism discussed how sensitive personal and company information within Anthropic’s Claude was exposed to Google when users used the “share” feature, making the information searchable to anyone on the web. While Claude notes that sharing will create a “public” link, which “anyone with the link can view”, it doesn’t make it clear to users that shared content could wind up being indexed by a search engine. The risks of exposing sensitive and confidential information within public LLMs are considerable.

By default (as of this writing), OpenAI’s ChatGPT, Anthropic’s Claude and Google Gemini train on your data (at least for personal accounts). This means when a party in litigation uploads a document produced by opposing counsel into any one of these public LLMs (or others), the model will train on that document, making it a possibility that information could be shared with an unauthorized party.

How the Public LLM Inadvertent Disclosure Risk Differs from the Use of Other Technology Platforms

Confidential information exchanged during litigation is increasingly exposed to the risk of inadvertent disclosure through the routine use of technology. Attorneys, legal support professionals, and corporate employees regularly rely on email, cloud-based collaboration platforms, file-sharing services, messaging applications, videoconferencing tools, and more to review and exchange information.

Common examples of inadvertent disclosures through technology include:

  • Emails sent to the wrong recipient because of auto-complete features
  • Documents uploaded to shared cloud repositories with overly broad permissions
  • Privileged or confidential files mistakenly attached to emails
  • Confidential information embedded in document metadata that is not removed before production

Even seemingly minor mistakes such as sharing an unrestricted hyperlink instead of a protected document or forwarding a message from a personal device can result in the unauthorized access of confidential litigation materials.

So, why is the use of public LLMs such a big deal when confidential information can be exposed via these other uses of technology? The examples above are mistakes that can happen. Best practices exist to minimise the risk of these mistakes occurring (e.g., procedures to follow before sending confidential documents in messages) and to mitigate the consequences if they do occur (e.g., follow up with unintended recipients to get them to delete the confidential information they received). Confidential documents are only disclosed by mistake and the extent of the disclosure is usually clear.

Conversely, when opponents use public LLMs to review opponents’ confidential documents, that is an intentional act that can put the opponent’s confidential information at risk, if the model is set to its normal default to train on user conversations. Not only that, but potential exposure of that information is virtually unlimited to all users of the model (and potentially beyond, given the shared chats example above) with no reasonable way to mitigate the risk.

Risk of Inadvertent Disclosure vs. Access to Justice

Given the risks discussed above, many legal professionals believe that restrictions should be put in place via protective orders to prohibit receiving parties from putting produced documents into a public LLM. Most have sought to protect sensitive and confidential documents, while some have even sought to protect any produced documents from being uploaded.

However, some legal professionals believe that these restrictions impede access to justice for parties that can’t afford a full-featured closed-loop eDiscovery AI solution. For example, in Craig Ball’s blog post titled The AI Protective Order Double Standard, he discusses how “all the legal AI products on the market run on the same small handful of foundation models including OpenAI’s GPT series, Anthropic’s Claude, Google’s Gemini and are hosted on the same cloud infrastructure, like Azure, AWS or Google Cloud.” Craig expresses his concern that not allowing parties to use lower cost public LLM tools will “widen the gap between well-funded litigants and everyone else, while delivering no meaningful improvement in data security.”

That’s an understandable concern. To be fair, owners of personal accounts can turn off the model’s default setting to train on their conversations. But many do not. Just as many lawyers and pro se parties don’t know enough about these tools to verify their outputs – leading to hallucinations in case filings – many don’t know enough about them to turn off the default for the model to train on their conversations. Firstly, their action creates risks for themselves in terms of embarrassment and sanctions by the courts. Secondly, their action creates risks, not for them, but for their opponents in terms of potential exposure of sensitive or confidential information. It’s understandable that parties want to limit the potential for unauthorized disclosure of their sensitive and confidential data.

Conclusion

Protecting the client’s sensitive and confidential information is one of the most important (if not the most important) responsibilities of lawyers and legal professionals. The ubiquity of public AI models and the tendency for their users to use them to analyse practically everything (including documents in litigation) increases the risk for exposure of that sensitive and confidential information considerably.

So, what are producing parties doing to help ensure that data is protected? Next time, we’ll discuss the use of protective orders to protect against use of public LLMs on produced data and provide some case law examples that illustrate how the courts are addressing these disputes.

For more regarding Cimplifi specialised expertise regarding AI & machine learning, click here.