7 Hidden Risks of Cloud AI for Law Firms and Business Data

A plain English guide to what actually happens to your prompts, your client files, and your financial records once you hand them to a cloud AI tool, and what a growing number of firms and business teams are doing about it.

A document and laptop send data upward into a cloud that leaks it back down into unmarked servers, illustrating how cloud AI tools expose confidential business and legal data
A document and laptop send data upward into a cloud that leaks it back down into unmarked servers, illustrating how cloud AI tools expose confidential business and legal data

Every legal assistant, finance manager, and operations lead who has pasted a contract, a client intake form, or a spreadsheet of vendor payments into a chatbot has run an experiment they never signed up for. The question was simple: can this tool save me an hour. The answer usually came back yes. What almost nobody stopped to ask was a second, quieter question: where does this information go once I hit enter, and who else can see it now.

That second question is the subject of this guide. It is not an argument against artificial intelligence. Teams that use AI well are faster, and firms that ignore it entirely will fall behind. The real issue is not whether you use AI. It is whether the version of AI you use sends your most sensitive information to a server you do not control, run by a company you have never audited, governed by a privacy policy nobody on your team has actually read.

For a mid size business team juggling financial forecasts, vendor contracts, and customer records, or a small legal practice holding privileged client files, that distinction is not academic. It is the difference between a productivity tool and a liability sitting quietly inside your workflow. Below, we walk through the risks that cloud based AI tools introduce, backed by court records, regulator filings, and the research firms that track this territory for a living, every claim linked to its original source so you can verify it yourself.

By the end, you will understand exactly what happens to a prompt after you send it, why courts and bar associations have already started punishing firms that got this wrong, and what a growing number of legal and business teams are doing instead: running AI locally, on hardware they own, where the data never leaves the building in the first place.

What “cloud based AI” actually means, and where your data really goes

It helps to be precise about the mechanics, because the phrase “cloud AI” hides a lot of plumbing that most users never see.

When you type a prompt into ChatGPT, Gemini, Copilot, Claude, or any other hosted assistant through a normal consumer or business account, that text does not stay on your device. It is packaged, encrypted in transit, and sent across the internet to a data center operated by the company behind the model, or increasingly, to a third party cloud platform that company rents capacity from. Once it lands there, several things can happen to it, and the honest answer is that as a user you usually do not know which ones apply to your specific message.

Your prompt may be logged and stored. Most providers keep a copy of your conversation on their servers for some period, ranging from a handful of days on a strict enterprise plan to indefinitely on a free consumer account, unless you take active steps to change that setting.

Your prompt may be used to train future models. On free and many entry level paid tiers, the default setting historically has been that conversations can be reviewed by human contractors and used to improve the underlying model, meaning fragments of your contract language, your client’s name, or your unreleased financial figures can, in principle, resurface in a future response to a completely different user. Enterprise agreements with contractual guarantees against training use exist, but they cost more, require a signed agreement, and are frequently skipped by teams that signed up with a personal credit card to save time.

Your prompt may be reviewed by human staff. Trust and safety teams at AI companies routinely sample conversations, both for quality control and to catch abuse, which means a real person, an employee of a company you have no contract with, may read what you typed.

Your prompt may sit inside a much longer supply chain than you think. The company whose logo is on the chat window rarely owns every server in the path. Cloud infrastructure is frequently subcontracted to a handful of hyperscale providers, and each of those providers has its own subprocessors, its own outsourced support staff, and its own list of jurisdictions where a copy of your data physically resides. A single prompt can legally and technically pass through several corporate entities and, depending on the provider, several countries, before it produces the answer you see on your screen.

Your prompt can outlive your delete button. This is the part almost nobody expects. Litigation, regulatory holds, and routine backup cycles can all keep a copy of a conversation alive long after you clicked delete. We will walk through a real, ongoing court case later in this guide where a federal judge ordered a major AI company to preserve hundreds of millions of user conversations, including ones users had explicitly deleted, because they might be relevant evidence in an unrelated lawsuit. If that can happen to a copyright dispute between a newspaper and an AI company, it can happen to a discovery request that touches your firm’s account too.

None of this makes cloud AI unusable. It makes it something you need to understand before you decide what belongs in it and what does not. The information a lawyer or a finance lead types into a prompt box is not private in the way an email to a colleague is private. It is more like handing a document to a photocopier operated by a stranger, in a building you have never visited, with a retention policy you have never read.

With that mechanical picture in place, here are the specific risks that follow from it, one at a time.

1. Your prompts can become someone else’s training data

Start with the risk that catches almost every new user off guard: the assumption that a chat window works like a private conversation. On many consumer AI products, it does not. Unless a user or an administrator has explicitly turned off model training and confirmed an enterprise agreement that excludes it, text typed into a prompt box can be stored, reviewed by human staff, and folded into the data used to train later versions of the model.

The clearest illustration of what this looks like in practice is not a hypothetical. In the spring of 2023, engineers at Samsung’s semiconductor division used ChatGPT to help with three separate work tasks: fixing faulty source code, optimizing a test sequence for identifying defective chips, and turning a recorded internal meeting into typed notes. In each case, they pasted genuinely confidential material, including proprietary source code, straight into the chat window. Samsung had no way to claw any of it back once it was sent, and within about three weeks of allowing the tool, the company had three documented leaks on its hands. Samsung’s response was blunt: it banned generative AI chatbots company wide and warned staff that violating the new policy could lead to termination.

What makes the Samsung case worth remembering years later is not the specific company. It is how ordinary the mistake was. These were not careless employees. They were skilled engineers trying to do their jobs faster, exactly the behavior every manager wants to encourage, using a tool their own employer had just approved weeks earlier. None of them set out to leak anything. They simply treated a chat window the way they would treat a search engine or a spellchecker, not realizing the destination on the other end was a server owned by an outside company with its own data practices.

That pattern has only grown since. Independent research from the security firm LayerX found that by 2025, 77 percent of enterprise employees using generative AI tools had copied and pasted data directly into a chatbot prompt, and that a striking 82 percent of those pastes came from personal accounts entirely outside company visibility or control, meaning the employer had no logging, no policy enforcement, and no way to know it was happening. Roughly a fifth of those copy and paste actions included personally identifiable or payment card information. Separately, the monitoring firm Cyberhaven, which analyzes what employees actually paste into AI tools across millions of prompts, found that the share of pasted corporate data classified as sensitive rose from 10.7 percent in 2023 to 34.8 percent in 2025, more than tripling in two years, according to figures reported by Unio Digital. A separate 2026 survey of five hundred employed adults in the United States, commissioned by the litigation firm Kolmogorov Law, found that 38 percent had entered at least one type of work information into a personal AI account their employer did not control, and that nearly two thirds did not know whether doing so could be against the law, a gap in awareness documented by outlets covering the survey.

For a business team, the exposure here is straightforward: unreleased financial projections, vendor pricing, customer lists, or a draft acquisition memo can all end up sitting on a server outside your control, subject to a privacy policy you never negotiated. For a legal team, the exposure is sharper, because the material being typed into that box is frequently privileged. A associate summarizing a client’s deposition, a paralegal asking an AI tool to shorten a settlement demand, or a partner drafting a sensitive email about litigation strategy is transmitting protected material to a third party the moment they hit enter, something we cover in detail in risk three below.

The fix is not complicated to state, even if it is hard to enforce: know, in writing, whether the AI tool your team uses trains on your input, how long it retains that input, and whether a human being can read it. If the answer to any of those questions is “we are not sure,” treat every prompt as if it were being published on a public forum, because functionally, that is closer to the truth than most people assume.

2. Cloud accounts are a single point of failure, and breaches are getting more expensive

Every cloud AI account your team creates is another login, another password, another set of servers holding your data, and another target. That matters because the economics of a data breach have shifted in a direction that should worry any small or mid size organization specifically.

According to IBM’s 2025 Cost of a Data Breach Report, the most cited independent benchmark in the industry, the global average cost of a data breach was 4.44 million dollars, and in the United States specifically that figure reached an all time high of 10.22 million dollars, roughly two and a half times the global average, driven by steeper regulatory fines and slower detection and escalation, according to IBM’s own analysis. The same report found something more specific to AI tools: among organizations that experienced a breach involving an AI model or application, 97 percent had no access controls in place for that AI system at all, and breaches involving unsanctioned “shadow AI” tools cost the organizations that had heavy shadow AI usage an extra 670,000 dollars on average, according to figures summarized by Northdoor’s analysis of the report.

Small and mid size organizations carry a disproportionate share of this risk, not because attackers value them more, but because they are easier to reach. Verizon’s 2025 Data Breach Investigations Report, drawing on more than 22,000 real world security incidents, found that small and mid size businesses experienced roughly four times as many confirmed breaches as large organizations, and separate industry compilations note that smaller firms receive the highest rate of targeted malicious email of any organization size, roughly one in every 323 messages, according to data aggregated by StrongDM and CNIC Solutions. Hiscox’s Cyber Readiness Report, surveying thousands of small and mid sized firms across multiple countries, has repeatedly found that a large share of small businesses hit by a serious cyber incident never fully recover their prior trajectory, and separate research from VikingCloud found that 40 percent of small business owners believe a cyberattack costing 100,000 dollars or less would be enough to put them out of business entirely.

Layer an AI vendor into that picture and you have added a large, high value, centrally stored pool of exactly the kind of data attackers want, sitting behind a login screen that is only as strong as the weakest password on your team. A breach at the AI vendor itself does not require your firm to make a single mistake. It only requires the vendor to make one, and you have no visibility into how well that vendor is actually defended.

Scales of justice beside a privileged document with a broken seal, connected by a dotted line to a cloud icon with an eye inside it, representing privileged legal information leaving a firm's control
Scales of justice beside a privileged document with a broken seal, connected by a dotted line to a cloud icon with an eye inside it, representing privileged legal information leaving a firm’s control

3. For lawyers, cloud AI creates a direct confidentiality problem, not just a security one

Business teams face a data protection problem when they use cloud AI carelessly. Law firms face something stricter: a professional duty they can be disciplined for breaching, independent of whether any breach or hack ever actually occurs.

Every lawyer in the United States operates under Model Rule of Professional Conduct 1.6, which protects, in the words of the American Bar Association’s own guidance, all information relating to the representation of a client, regardless of its source, unless the client consents to disclosure. In July 2024, the ABA’s Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, its first formal guidance addressing generative AI directly, and the core instruction is unambiguous: before a lawyer inputs any information relating to a client’s representation into a generative AI tool, the lawyer must evaluate the risk that the information could be disclosed to or accessed by people outside the firm, according to the opinion summarized by Zuva and the National Conference of Bar Examiners. The opinion draws a direct line to the ABA’s earlier guidance on cloud computing and outsourcing, treating a consumer AI chatbot the same way it would treat handing client files to an unvetted outside vendor.

This is not a hypothetical compliance exercise. It is a live, escalating body of case law and disciplinary action, and state bars have followed the ABA’s lead with their own opinions covering everything from client consent requirements to fee billing for AI assisted work. A lawyer who pastes a client’s financial records, a settlement number, or deposition testimony into a general purpose chatbot to save drafting time may be violating Rule 1.6 the moment that text leaves the firm’s systems, regardless of whether OpenAI, Google, or any other vendor ever suffers a breach. The duty is triggered by loss of control over the information, not by proof of misuse.

Several major law firms and financial institutions recognized this early and simply blocked consumer AI tools outright. J.P. Morgan Chase and Verizon restricted employee access to ChatGPT in 2023 specifically over data handling concerns, according to reporting from Gizmodo, and law firm technology surveys since have repeatedly found that firms are far more willing to adopt AI for internal drafting and research than for anything touching client identifying detail, precisely because the confidentiality calculus is so much less forgiving in legal practice than in a typical business context.

4. AI hallucinations are landing lawyers in front of judges, in growing numbers

A large language model does not know facts the way a database knows facts. It predicts the next plausible word based on patterns in its training data, which means it can produce a case citation that reads exactly like a real one, complete with a plausible docket number, a real judge’s name, and confident sounding legal reasoning, while describing a case that never existed. Researchers call this a hallucination. Judges increasingly just call it a problem for the lawyer who filed it.

The case that put this on the map was Mata v. Avianca, decided in the Southern District of New York in June 2023. Attorneys representing a plaintiff in a personal injury claim submitted a brief opposing the airline’s motion to dismiss, citing six court decisions as supporting precedent. All six were fabricated by ChatGPT. When opposing counsel could not locate the cases, the presiding judge, P. Kevin Castel, ordered the plaintiff’s lawyers to produce copies. What followed was worse than the original error: rather than admitting the mistake immediately, the attorney who had used the tool submitted an affidavit containing purported excerpts from the fake decisions, themselves also generated by ChatGPT, according to the detailed case record compiled by LegalClarity. One attorney had even asked the chatbot directly whether the cases were real, and it confidently told him yes. Judge Castel ultimately fined the attorneys and their firm 5,000 dollars and ordered them to send a personal letter of correction to every judge whose name had been falsely attached to a fabricated opinion, a detail confirmed in the Wikipedia summary of the underlying court order.

At the time, Mata read like a strange, isolated story. It was not. Legal researcher Damien Charlotin, a fellow at HEC Paris, maintains the most comprehensive public tracker of these incidents, and the trajectory is worth sitting with: his database held roughly 200 documented cases of AI hallucinated material in court filings in mid 2025, 719 by January 2026, over 1,200 by early April 2026, and more than 1,600 by June 2026, according to reporting compiled by HAQQ and PlatinumIDS, which described the pace as roughly five to six newly documented cases every single day.

The penalties have escalated alongside the volume. Where Mata resulted in a 5,000 dollar sanction, later cases have gone much further. In one Oregon commercial dispute, attorneys who filed briefs containing fifteen fabricated citations and eight invented quotations faced combined sanctions, fines, and opposing counsel’s fees totaling roughly 109,700 dollars, described by the presiding magistrate as a notorious outlier in scale, according to HAQQ’s tracker. A Nebraska attorney was placed under an interim suspension from practicing law in 2026 after filing a divorce appeal in which 57 of 63 citations were defective, including twenty entirely invented cases, the first reported instance of a US bar removing a lawyer from practice specifically over AI generated filings. Courts in Singapore and the United Kingdom have gone further still, holding supervising partners personally liable for failing to check a junior associate’s AI generated work, explicitly rejecting heavy workload as a defense.

None of this requires malice. Every documented case involves a professional who believed the tool’s output was accurate enough to file. That is precisely why it matters for your firm’s evaluation of any AI tool: the question is not whether an AI assistant can save time on research or drafting, it clearly can, but whether the specific tool your team relies on is grounded in a verified, closed set of documents you control, versus a general purpose model pulling from the open internet with no way for it, or you, to know when it is confidently wrong. A verification step that catches a fabricated citation before it reaches a judge is no longer optional diligence. Formal Opinion 512 makes clear that the duty to check AI output sits with the lawyer who signs the filing, not with the tool that generated it.

5. Your deleted chats may not actually be deleted

Most people assume that clicking delete on a conversation ends its life. A federal court case working through 2025 and into 2026 shows why that assumption can be wrong, and why the implications reach well beyond the two companies actually in the courtroom.

In its ongoing copyright litigation against OpenAI, The New York Times argued that some ChatGPT users were using the tool to get around the newspaper’s paywall, and that conversations relevant to proving this might be the ones users were most likely to mark as temporary or delete. On May 13, 2025, a magistrate judge sided with the Times and ordered OpenAI to preserve and segregate all output log data that would otherwise be deleted, on a going forward basis, for the entire active consumer user base across free, Plus, Pro, and Team tiers, according to the order described by the Center for Democracy and Technology. In practice, that meant conversations users believed were gone, including ones covering health questions, business plans, and personal matters entirely unrelated to the Times lawsuit, were instead being kept on file as potential evidence in someone else’s case.

OpenAI fought the order and eventually succeeded in narrowing it. According to OpenAI’s own account of the litigation, the company’s obligation to indefinitely retain new consumer content ended on September 26, 2025, and the company returned to standard practice, deleting most conversations within 30 days. But the episode left a permanent lesson rather than a temporary one: an AI provider’s retention promises are only as durable as the next lawsuit a court decides is relevant enough to override them. Business customers using a specific contractual arrangement called zero data retention were unaffected throughout, because their prompts were never stored on OpenAI’s servers to begin with, which tells you exactly what the safer alternative looks like, and how few default consumer or entry level accounts actually have it.

For a law firm, the stakes compound further. If your firm’s own AI usage becomes relevant to a malpractice claim, a bar complaint, or a discovery dispute in a client’s litigation, prompts and outputs stored on a vendor’s servers, not yours, are now treated by courts as discoverable material subject to the same preservation and production rules as email and Slack messages, according to the analysis published by Terms.law. You do not control when that data is produced, to whom, or under what protective order, because you are not the party being subpoenaed. Your AI vendor is. And your firm’s confidential material is sitting inside the response.

6. Cross border rules and regulatory exposure are catching up fast, and unevenly

Where your data physically sits, and which regulator has authority over it, used to be a question only multinational corporations worried about. Cloud AI has made it a question for a five person legal practice too, because most consumer and mid tier AI accounts route data through servers and subprocessors located wherever the vendor’s infrastructure happens to be, often without the customer being told exactly where.

The clearest cautionary tale here is Italy’s regulator, the Garante, which in March 2023 became the first authority in the world to order a temporary block of ChatGPT within its borders, citing the absence of a lawful basis for processing users’ personal data to train the model, missing age verification for minors, and a failure to properly disclose a March 2023 data breach that had briefly exposed some users’ chat titles and payment details to other users. The investigation ran for nearly two years and concluded in December 2024 with a 15 million euro fine against OpenAI, along with an order to run a six month public information campaign about how the tool handles personal data, according to reporting from The Hacker News and legal analysis from the National Law Review. It was the first fine of its kind against a generative AI company, and other European data protection authorities, including Germany, France, and Spain, opened their own parallel investigations in the same period, coordinated through a dedicated European Data Protection Board task force on ChatGPT.

Since then, the regulatory floor has kept rising rather than settling. The European Union’s AI Act, the first comprehensive AI specific law of its kind, entered into force in August 2024 and rolled out obligations in stages, with the heaviest requirements for high risk AI systems, covering categories like employment decisions, credit scoring, and access to essential services, originally scheduled to bind operators by August 2026. In 2026 EU lawmakers agreed to push most of those high risk deadlines back to December 2027 to give companies more runway, a deferral confirmed by Morgan Lewis and Gibson Dunn, but the transparency and general purpose AI obligations under the Act still took effect on schedule in August 2026, and penalties for noncompliance with the high risk rules once they land can reach 15 million euros or 3 percent of a company’s global annual turnover, whichever is higher, according to the Cloud Security Alliance’s research summary.

In the United States, there is no single federal equivalent, which creates its own hazard: a patchwork of state privacy laws, sector rules like HIPAA for health information, and state bar ethics opinions that do not always agree with each other on what a lawyer or a business owes clients when using AI. A business team assuming that a cloud AI vendor’s headline privacy policy covers them in every state or country they operate in is often assuming more legal coverage than actually exists. A vendor’s marketing page describing itself as compliant is not the same thing as a signed agreement specifying where your data is processed, stored, and under whose jurisdiction it sits. Reading the actual data processing addendum, not the summary page, is the only way to know which regulator, if any, would have jurisdiction over a breach involving your firm’s information.

7. You are locked into someone else’s roadmap, pricing, and risk tolerance

The final risk is the least dramatic and the most certain to affect every team that adopts a cloud AI tool: you do not control the vendor’s decisions, and every one of those decisions becomes your problem the moment your workflow depends on it.

Cloud AI vendors change data handling policies, sometimes without much warning. They deprecate older models your prompts and internal tooling were tuned around, sometimes with a matter of weeks notice. They renegotiate enterprise pricing at renewal, often steeply, once a team is dependent enough that switching costs feel prohibitive. They suffer outages, during which a firm that has built document review, client intake, or financial reporting workflows around a live API connection simply stops working until the vendor’s engineers fix the problem, on their timeline, not yours. And in the event of an acquisition, a bankruptcy, or a strategic pivot away from the product your business relies on, your firm inherits that disruption with essentially no leverage, because you were a customer, not a partner, in the arrangement.

There is a subtler version of this risk too, one that is easy to miss because it never shows up as an outage or a price increase: your competitive intelligence becomes part of an aggregate pool you cannot see or audit. A cloud AI vendor serving thousands of businesses in the same industry sees patterns across all of them, the kinds of questions being asked, the kinds of language showing up in draft contracts, the shape of pricing strategies being tested. Reputable vendors build safeguards against any single customer’s data leaking into another customer’s output, but those safeguards are a promise, not a physical impossibility, and a small or mid size firm has no practical way to audit a hyperscale vendor’s internal controls to confirm the promise is being kept. Running the model yourself, on your own infrastructure, removes the question entirely, because there is no aggregate pool for your data to fall into in the first place.

None of these seven risks require a hacker, a nation state, or a dramatic villain. They are the ordinary, foreseeable consequences of sending sensitive information to infrastructure you do not own, run by a company whose incentives are not identical to yours. That is worth sitting with before the next section, which looks at how these risks actually played out in the real world, not as abstractions, but as dated, named, verifiable incidents.

What this means specifically for finance and operations teams

Most of the incidents and rules covered so far center on legal practice, because the professional duties involved are unusually well documented and the disciplinary consequences unusually visible. Business and finance teams face a version of the same exposure, quieter but no less real, and worth naming specifically rather than assuming the legal examples translate automatically.

Unreleased financial figures are exactly the kind of material employees paste into AI tools without thinking twice. A finance manager asking a chatbot to help phrase a board memo, summarize a quarterly forecast, or clean up a spreadsheet of vendor payment terms is transmitting information that, if it reached a competitor, an activist investor, or the open market ahead of a public announcement, could carry real financial and legal consequences. The Kolmogorov Law survey referenced earlier in this guide found that more than one in ten workers admitted specifically to entering financial or sales figures into a personal AI account outside employer control, a detail reported by Alabama Gazette’s coverage of the survey. Unlike a legal citation, a leaked financial figure does not need a judge to notice it for the damage to be real.

Vendor and pricing data is competitive intelligence, whether or not your team thinks of it that way. Contract terms, negotiated discounts, supplier lists, and margin structures are frequently the exact material an operations or procurement lead pastes into an AI tool while drafting a renewal negotiation or comparing supplier proposals. If that data becomes part of an aggregate pool a cloud vendor cannot fully guarantee is siloed between customers, and if a competitor in the same industry uses the same popular AI tool, the theoretical risk of cross contamination described earlier in this guide stops being theoretical for the business it actually happens to.

Mergers, acquisitions, and other market moving activity carry a sharper version of every risk in this guide. A due diligence summary, a draft term sheet, or an internal valuation model is precisely the kind of document where a data leak, a discoverable chat log, or an unverified AI hallucination in a financial model has consequences measured in real money, not just reputational discomfort. Deal teams handling this kind of material benefit disproportionately from a local, fully contained AI setup, because the sensitivity of the material is at its peak during exactly the period when the widest possible group of people, bankers, lawyers, auditors, and executives, are all working with the same documents under time pressure.

Customer and payment data brings its own regulatory layer on top of everything already covered. Depending on what a business collects, additional frameworks like state data breach notification laws, payment card industry standards, and sector specific rules can all apply on top of the general privacy and AI governance questions covered throughout this guide. A business that has never mapped which of its AI tools touch customer records that fall under these frameworks is carrying regulatory exposure it has not yet quantified, and IBM’s finding that breaches involving unsanctioned shadow AI tools cost affected organizations an average of 670,000 dollars more than other breaches suggests that exposure is not a rounding error.

The throughline for finance and operations teams is the same one running through the legal sections of this guide: the risk is rarely a single dramatic event. It is the accumulation of small, individually reasonable feeling decisions, an analyst pasting a spreadsheet here, a manager summarizing a memo there, none of which feel risky in the moment, adding up to an exposure nobody specifically chose to create. The fix looks the same too: an honest inventory of where sensitive financial and vendor data currently goes, a clear policy about what belongs in a cloud tool versus what does not, and, for the material that genuinely cannot afford exposure, a local AI setup that removes the question of trust entirely.

A timeline of what has actually happened, not what might happen

It is easy to treat data privacy risk as theoretical until you see how often it has already materialized. Below is a condensed, dated record of the incidents referenced throughout this guide, along with a few others worth knowing, so you can see the pattern rather than a single anecdote.

March 2023. Italy’s data protection authority, the Garante, issues an emergency order temporarily blocking ChatGPT across the country, triggered by a data breach that briefly exposed some users’ chat histories and payment details to other users, and by concerns over the lack of a lawful basis for using personal data to train the model. Access is restored about a month later after OpenAI adds a privacy notice, an opt out for training data, and age verification. Source: Data Protection Report.

March to April 2023. Samsung engineers paste confidential source code, defect optimization data, and internal meeting notes into ChatGPT on three separate occasions within about twenty days. Samsung bans generative AI chatbots company wide within weeks. Source: Forbes.

June 2023. A federal judge in the Southern District of New York sanctions two attorneys 5,000 dollars in Mata v. Avianca after they file a brief containing six court decisions fabricated by ChatGPT, and orders them to notify every judge falsely named in the fake opinions. Source: Wikipedia case summary.

July 2024. The American Bar Association issues Formal Opinion 512, its first formal ethics guidance covering lawyers’ use of generative AI, requiring lawyers to evaluate confidentiality risk before inputting any client related information into an AI tool. Source: American Bar Association.

December 2024. Italy’s Garante fines OpenAI 15 million euros, the first GDPR penalty issued anywhere against a generative AI company, citing an inadequate legal basis for training on personal data and a failure to properly report the March 2023 breach. Source: The Hacker News.

May 2025. A magistrate judge orders OpenAI to preserve all ChatGPT output logs that would otherwise be deleted, including conversations users had explicitly deleted, as part of discovery in The New York Times’ copyright lawsuit. The order is later narrowed, and OpenAI’s obligation to retain new consumer data indefinitely ends in September 2025. Source: Center for Democracy and Technology.

Throughout 2025 and into 2026. The publicly tracked count of court cases involving AI hallucinated legal citations climbs from roughly 200 in mid 2025 to more than 1,600 by June 2026, spanning multiple countries, with penalties escalating from a five figure fine in the earliest cases to six figure sanctions and, in at least one instance, an interim license suspension. Source: HAQQ sanctions tracker.

2025 to 2026. Independent research firms including IBM, LayerX, and Cyberhaven document a sharp rise in employees pasting sensitive company data into AI tools outside employer control, with LayerX finding 77 percent of AI using employees had copied data into a prompt and IBM finding that 97 percent of organizations breached through an AI system had no access controls on that system at all. Sources: The Register and Help Net Security.

What this timeline shows is not a string of freak accidents. It shows the same handful of failure modes repeating across different companies, different countries, and different regulators: data leaving a firm’s control without anyone deciding it should, AI output being trusted without verification, and legal or regulatory processes reaching into stored conversations long after the people who typed them assumed the matter was closed.

Why small legal firms and mid size business teams carry the most exposure, not the least

There is a comforting myth that cyber and AI related risk is mainly a large enterprise problem, the kind of thing that happens to household name companies with billions of records at stake. The data says the opposite. Smaller organizations are targeted more often, recover more slowly, and have far less room to absorb the cost of getting this wrong.

Start with the plain economics. Verizon’s 2025 Data Breach Investigations Report found that small and mid size businesses experience roughly four times as many confirmed breaches as large organizations, and receive the highest rate of targeted malicious email of any organization size category, according to figures compiled by CNIC Solutions. Nearly half of businesses with fewer than 50 employees allocate no dedicated cybersecurity budget at all, and only a small minority of small businesses carry cyber liability insurance, according to the same aggregation. A large enterprise absorbing a multi million dollar breach dents a quarterly earnings report. A twelve person law firm or a forty person business team absorbing even a fraction of that cost can lose the practice entirely, and VikingCloud’s 2025 research found that a striking share of small business owners believe an incident costing as little as 100,000 dollars would be enough to end their business.

Legal practices carry an additional layer most business teams do not. The American Bar Association’s 2025 TechReport found that 29 percent of law firms have experienced a security breach at some point, with firms in the ten to forty nine attorney range reporting the highest incident rates of any firm size category, according to analysis from Petronella Cybersecurity. Among firms that experienced a breach, over half reported losing sensitive client information specifically, according to figures reported by Embroker. Larger firms typically employ dedicated IT security staff, run regular penetration testing, and can absorb the fixed cost of a full compliance program across hundreds of partners. A small firm is splitting that same fixed cost, and that same exposure, across a handful of people, often without a single staff member whose full time job is security.

There is a second, less obvious reason small teams are exposed: they are the ones most likely to adopt cloud AI tools through informal, personal channels rather than a vetted enterprise agreement, precisely because a formal procurement process feels like overkill for a five or ten person team. A large corporation typically routes new software through legal, procurement, and IT security review before anyone is allowed to type client data into it. A small firm’s version of that process is often a partner or office manager signing up for a monthly subscription with a company card on a Tuesday afternoon because a webinar made it look useful. The review that would have caught the training data setting, the retention window, or the data residency question simply never happens, not because anyone was careless, but because nobody on a lean team has the bandwidth to be the dedicated skeptic.

None of this means small firms and mid size teams should avoid AI. It means the tools they choose, and the way those tools are configured, deserve more scrutiny than their size and resources typically allow them to give, which is exactly the gap that keeping AI processing local and self contained is designed to close.

What regulators, courts, and bar associations are actually saying right now

Strip away the acronyms and the pattern across every regulator, court, and professional body watching this space is consistent: the burden of proof has shifted onto the organization using the AI tool, not the vendor that built it.

In the United States, the National Institute of Standards and Technology published its AI Risk Management Framework in January 2023, the first comprehensive federal guidance on managing AI risk. It is voluntary, but it has become the reference point courts, regulators, and cyber insurers increasingly measure organizations against when a dispute arises, because it lays out, in plain terms, the categories of risk, including data privacy, that any organization deploying AI is expected to have actually thought about, not merely hoped would work out. Internationally, ISO/IEC 42001, published in December 2023, became the first management system standard specifically for AI, giving organizations a certifiable framework similar in spirit to the ISO 27001 information security standard many businesses already know.

For lawyers, the guidance has moved from general to specific with unusual speed. Formal Opinion 512 remains the anchor, but state bars have layered their own opinions on top of it, several of them going further than the ABA baseline on issues like informed client consent before AI use, disclosure of AI assisted work in fee billing, and the standard of review expected before an AI drafted document reaches a client or a court. The throughline across nearly all of them, regardless of state, is that the confidentiality duty does not pause because a task got outsourced to software. A lawyer cannot delegate away Rule 1.6 any more than they could hand a client’s file to an unvetted temp worker and call the confidentiality question solved.

Courts, for their part, have stopped treating AI hallucinations as a novel curiosity deserving leniency. The earliest cases, including Mata v. Avianca, drew real but comparatively modest sanctions, and judges were often willing to note that nothing is inherently improper about using AI as a tool. By 2026, that patience has visibly thinned. Sanctions orders increasingly describe unverified AI citations as a straightforward violation of the basic professional duty every filing lawyer already owed the court, long before generative AI existed, to personally verify what they submit. Judges in multiple jurisdictions have begun holding supervising partners liable for failing to check a subordinate’s AI assisted work, rejecting workload and time pressure as an excuse.

On the regulatory side, the direction of travel is the same everywhere, even where the specific rules differ. The EU AI Act, despite its delayed high risk timeline, still brought binding transparency and general purpose AI obligations into force in August 2026, and its penalty structure, up to 15 million euros or 3 percent of global turnover, sets a ceiling few businesses can shrug off. In the United States, the absence of a single comprehensive federal AI privacy law has not meant an absence of enforcement risk. It has meant a denser, less predictable patchwork of state privacy statutes, sector specific rules like HIPAA for anything touching health information, and state attorney general actions, each with its own definitions and thresholds a business has to track individually rather than through one unified compliance checklist.

None of these frameworks, taken individually, tell a business exactly what to do. Taken together, they tell you what regulators, courts, and bar associations all agree on: that using a third party AI tool does not transfer responsibility for the data inside it, and that the organizations best positioned heading into 2027 are the ones that can answer, specifically and in writing, where their data goes, who can access it, and how long it is kept, for every AI tool in active use. Very few cloud AI setups, especially the ones adopted informally through a personal account or a quick signup, can currently answer that question with confidence. That gap is exactly what is driving the shift toward processing AI locally, which the next section covers in detail.

A desktop computer protected by a shield and closed padlock, with data circulating in a closed loop around it while a separate, faded cloud icon sits disconnected and crossed out in the distance
A desktop computer protected by a shield and closed padlock, with data circulating in a closed loop around it while a separate, faded cloud icon sits disconnected and crossed out in the distance

Why running AI locally closes most of these gaps at the source

Every risk covered so far shares one root cause: your data has to leave your building, and pass through infrastructure you do not control, for a cloud AI tool to work at all. Local AI removes that requirement entirely. Instead of sending a prompt across the internet to a remote data center, the model runs directly on hardware you own, typically a capable desktop, workstation, or small dedicated server sitting in your office. The prompt, the document, the client file, the spreadsheet, all of it stays on a machine you control, is never transmitted to a third party, and cannot be logged, reviewed, subpoenaed from a vendor, or folded into someone else’s training data, because there is no vendor server in the path to begin with.

Walk back through the seven risks above and the pattern is direct rather than partial. There is no training data risk, because a locally run model processing your documents is not phoning home with them. There is no single point of failure sitting on a vendor’s servers waiting to be breached, because there is no centralized pool of your firm’s data anywhere outside your own network to breach. The confidentiality duty under Rule 1.6 becomes dramatically easier to satisfy, because the information relating to a client’s representation never leaves the systems the firm already controls and has already secured. Hallucination risk does not disappear, because it is a property of how language models generate text, not of where they run, but a properly configured local system can be paired with your firm’s own verified documents and a mandatory review step far more tightly than a general purpose cloud chatbot most employees access on a personal account. There is no discoverable vendor held chat history to be preserved under a court order in an unrelated lawsuit, because there is no vendor holding it. Cross border and data residency questions resolve themselves, because the data was never routed anywhere outside the country, let alone the building. And there is no vendor roadmap risk, no surprise pricing change, no sudden deprecation, no aggregate pool of competitive intelligence you cannot audit, because you own the hardware, and typically the model weights, outright.

This is not a fringe idea anymore, and it is worth being precise about why. Model quality that used to require a data center now runs credibly on a well specified desktop or workstation. Industry analysts tracking this space describe a clear three stage adoption curve: capable local models becoming a normal tool for developers and privacy conscious professionals through 2026, private, self hosted AI infrastructure becoming a serious option for small teams and mid size businesses by 2027 as hardware costs continue falling and managed local deployment tools mature, and broad affordability for small businesses following by 2028, according to the analysis published by PromptQuorum. The underlying market reflects the same trajectory from the hardware side. The global edge AI market, covering AI processing that happens on local devices rather than centralized cloud servers, was valued at roughly 24.9 billion dollars in 2025 and is projected by Grand View Research to reach 118.7 billion dollars by 2033, a compound annual growth rate above 21 percent, with the analysis specifically citing data privacy and reduced dependence on centralized cloud infrastructure as core drivers of enterprise adoption, according to Grand View Research’s market report.

There is an honest tradeoff worth naming rather than glossing over. The very largest cloud hosted frontier models still edge out locally deployable open models on some general benchmark comparisons, though industry researchers tracking the gap describe it as narrowing steadily year over year rather than static, according to the PromptQuorum analysis cited above. For a huge share of the actual daily work a legal or business team needs from AI, drafting, summarizing, reviewing contracts against a firm’s own precedent library, searching a firm’s own document set, answering questions grounded in a company’s own financial data, that gap rarely matters, because the task is bounded and the model is working from your own material rather than trying to recall the entire internet from memory. What you gain in exchange is control that no contractual promise from a cloud vendor can fully replace: the physical, verifiable fact that your most sensitive information never left the room.

This is precisely the gap a new generation of privacy first, agentic software is being built to close, tools designed from the ground up to run entirely on a PC you already own, with no cloud dependency, no training data question, and no vendor to subpoena, because there is no vendor sitting between your team and your own files. If your firm or your business handles the kind of material covered in this guide, privileged client information, unreleased financial data, competitively sensitive strategy, it is worth watching this category closely over the next year, and testing a local option before your next cloud AI renewal locks you in for another twelve months.

A practical checklist for evaluating any AI tool before it touches client or company data

Whether your team ultimately chooses a cloud platform with strong contractual protections, a hybrid setup, or a fully local deployment, the questions below are the ones worth answering in writing before a single client file or financial record goes anywhere near the tool. Treat a vendor’s inability, or reluctance, to answer any of these clearly as an answer in itself.

Where is the data processed, and where does it live afterward. Ask the vendor to name the specific country or countries where processing happens and where any stored copies sit, not the region described in marketing copy. A vague answer here is disqualifying for anything touching client or financial data, because you cannot evaluate a cross border compliance question you cannot get a straight answer to.

Is your input used to train the model, by default or otherwise. Many providers offer an enterprise tier where training use is contractually excluded, while the free or entry level consumer tier of the exact same product trains by default. Confirm which tier your team is actually using, not which tier the vendor’s homepage advertises, and get the training exclusion in writing as part of a signed agreement rather than a checkbox in account settings that could be reset in a future update.

What is the retention window, and does it survive deletion. Ask specifically what happens when a user deletes a conversation. Is it gone immediately, held in backups for a defined period, or potentially preserved indefinitely under circumstances like litigation, as happened in the OpenAI case described earlier in this guide. A vendor that cannot describe its own deletion mechanics in specific terms has not built one worth trusting.

Who can access your data internally, and under what conditions. Human review for quality control and abuse prevention is standard across the industry, but the scope, frequency, and staffing of that review varies enormously between vendors. Ask whether review staff are direct employees or an outsourced contractor, and whether access requires a specific business justification and audit log, or is broadly available across the vendor’s support organization.

What is the full list of subprocessors, and are they disclosed proactively. A modern AI product is rarely a single company’s infrastructure end to end. Ask for the current subprocessor list, which should be a standard, available document for any vendor serious about enterprise sales, and check whether the vendor commits to notifying you before adding a new one, rather than updating a webpage you are expected to monitor yourself.

Does the contract include a real data processing agreement, not just a privacy policy. A privacy policy is a unilateral statement the vendor can often change at will. A signed data processing agreement is a binding contractual document, typically required under GDPR and similar frameworks for any relationship where the vendor is processing personal data on your behalf. If a vendor cannot produce one, treat that as equivalent to no contractual privacy protection existing at all.

What happens to your data, and your access to your own history, if you cancel. Ask how quickly your data is deleted after termination, whether you can export your own conversation and document history first, and whether the vendor has any right to retain a copy after the relationship ends. Firms that never ask this question tend to find out the answer at the worst possible moment, mid dispute with a former vendor.

Can the tool operate without an internet connection, or does every request necessarily leave the building. This is the single question that separates a genuinely local AI deployment from a cloud tool with a local sounding name. If a product requires an active internet connection for every single request, your data is going somewhere outside your network for that request, regardless of how the marketing describes it. A true local deployment should function, and should be verifiable as functioning, with the network cable unplugged.

Who owns the model weights and the infrastructure the tool runs on. In a cloud arrangement, the vendor owns both, and your access is a subscription that can be modified, repriced, or revoked. In a local deployment, your firm owns the hardware and typically the model files themselves, which removes the vendor lock in risk described earlier in this guide almost entirely.

Running through these ten questions with any vendor, cloud or local, tends to surface the honest answer quickly. Vendors built around a genuine privacy first architecture answer most of them in a sentence each, because the answers are structurally simple: your data, your hardware, your control. Vendors built around a cloud first architecture with privacy features layered on top tend to answer with more caveats, more tiers, and more references to policies that can change.

Building a simple AI use policy your team will actually follow

A written AI policy does not need to be long to be effective. The firms and business teams that manage this risk well tend to share a small number of common elements, adapted to their size, rather than an elaborate governance framework borrowed from a much larger organization.

Name the categories of data that require extra care, in specific terms. Rather than a vague instruction to be careful with sensitive information, list what actually counts: client identifying details, privileged communications, unreleased financial figures, vendor and pricing terms, employee personal data, anything covered by a signed confidentiality agreement. A specific list is something a busy employee can actually check a document against in ten seconds. A vague instruction is not.

Draw a clear line between approved and unapproved tools. Every team has a tool it has vetted, configured correctly, and trusts for a defined set of tasks, and a much longer list of tools employees have simply never been told not to use. Publish the approved list, explain briefly why each unapproved tool did not make the cut, whether that is training data defaults, unclear data residency, or simply not having been reviewed yet, and make it easy for someone to request review of a new tool rather than quietly adopting it on a personal account.

Require verification of anything AI produces before it reaches a client, a court, or a financial decision. Given the hallucination pattern documented throughout this guide, a policy that stops at “use AI responsibly” is not specific enough to hold up under scrutiny after something goes wrong. Spell out exactly what verification looks like for your practice, checking every citation against a primary source before filing, confirming every financial figure against the underlying system of record, having a second person review anything client facing that AI helped draft.

Put someone specific in charge of the policy, and give them the authority to say no. In a small firm or team, this does not need to be a dedicated role. It needs to be a named person, a managing partner, an operations lead, whoever is best positioned, who owns keeping the approved tool list current, reviewing new requests, and fielding questions when someone is unsure whether something is allowed. A policy with no clear owner tends to quietly stop being followed within a few months.

Review the policy on a fixed schedule, not only after something goes wrong. This space has moved quickly enough that a policy written even a year ago may already be missing tools, regulations, or bar guidance that did not exist when it was drafted. A short, scheduled review, quarterly for a fast moving practice, at minimum annually for anyone, keeps the policy matched to the actual tools your team is using, rather than the tools it was using when the document was first written.

None of this needs to slow a team down. A policy that takes fifteen minutes to read and is actually followed protects a firm more than an exhaustive document nobody has time to finish.

Honest answers to the objections local AI usually gets

Every team that gets this far in the conversation raises the same handful of concerns. They deserve straight answers rather than a sales pitch, because a genuinely good decision here should survive scrutiny.

“Won’t a cloud model just be smarter.” On raw, general knowledge benchmarks, the very largest cloud hosted models still have an edge over what most teams can run locally today, and that gap is real. But most of the work a legal or business team actually needs from AI is not a test of general world knowledge. It is drafting from a template, summarizing a document already in front of the model, comparing a contract clause against a firm’s own precedent library, or answering a question about a company’s own financial data. For grounded, bounded tasks like these, a well configured local model working from your own material closes most of the practical gap, and the portion of the gap that remains keeps shrinking every year as smaller models improve.

“Setting this up sounds like a serious IT project.” It used to be. The current generation of local AI tools, including the interfaces and frameworks the open source community has built around them, has moved this from a specialized infrastructure project to something closer to installing a well designed desktop application. A team does not need an in house engineer to get a capable local setup running on a single well specified machine, and dedicated privacy first products built specifically for legal and business use are designed to remove even the modest technical steps that remain.

“What about updates, if there’s no cloud team pushing new versions.” Local does not mean frozen. Open and locally deployable models are updated on an active, ongoing basis by the research community and by vendors building products on top of them, and a well designed local tool lets you pull in a new model version on your own schedule, after your team has had a chance to evaluate it, rather than having a cloud vendor silently swap the model underneath you overnight, which happens more often than most users realize.

“Our team is small. Can we really maintain this ourselves.” This is precisely the argument for a dedicated, purpose built product rather than a raw open source stack assembled from scratch. A raw self hosted setup does ask more of a small IT footprint than most lean teams have to spare. A local AI product built specifically for legal and business teams is designed to remove that burden, packaging the model, the interface, and the security configuration into something a small team can run without a dedicated engineer, the same way a modern accounting or practice management platform does not require a database administrator on staff.

“Don’t we lose collaboration features cloud tools offer.” Multi user access, shared workspaces, and permission controls are not exclusive to cloud architecture. They can run on a local server sitting on your own office network just as easily as they run on a vendor’s cloud servers, with the difference being that the server, and every byte of data on it, belongs to you rather than to a company billing you monthly for the privilege of storing your own files.

“Isn’t the hardware expensive.” It was a bigger barrier two or three years ago than it is now. Consumer and small business grade hardware capable of running genuinely useful local models has become significantly more affordable and more widely available, part of why the edge AI hardware market itself is projected to keep growing at a double digit rate through the rest of the decade. For many teams, the hardware cost compares favorably against a year or two of a growing monthly per seat cloud AI subscription, and unlike a subscription, the hardware is an asset your firm still owns after the first year.

Frequently asked questions

Is it safe to use ChatGPT or similar tools for legal documents.
It depends entirely on the account type and settings. A free or entry level consumer account typically does not include a contractual guarantee against training use, and courts have already shown, in the OpenAI preservation order described above, that even deleted conversations can be recovered under legal process. An enterprise agreement with a signed data processing addendum and confirmed training exclusion is meaningfully safer, but even then, the American Bar Association’s Formal Opinion 512 requires a lawyer to independently evaluate the confidentiality risk before inputting any client related information, regardless of tier.

Can law firms use AI at all under current ethics rules.
Yes. No bar association or ethics opinion, including Formal Opinion 512, prohibits AI use outright. The obligation is to use it competently, verify its output, protect client confidentiality in the process, and disclose its use to clients where relevant, not to avoid the technology entirely.

What does shadow AI mean, and why does it matter.
Shadow AI refers to AI tools employees use for work without their employer’s knowledge, approval, or configuration, typically through a free personal account. It matters because it removes an organization’s ability to enforce data handling policies, log usage, or even know that sensitive information is being shared with an outside vendor in the first place. IBM’s 2025 research found that unsanctioned AI use was a contributing factor in a meaningful share of the breaches it studied.

Does ChatGPT or similar tools use my prompts to train future models.
It depends on the product and the settings. Free consumer tiers have historically defaulted to allowing training use unless a user manually opts out. Enterprise and API accounts, particularly those with a zero data retention agreement, generally exclude training use by contract. The only reliable way to know for a specific account is to check the current settings and, for business use, the signed agreement rather than the general marketing page.

Can deleted AI conversations really be used as evidence in a lawsuit.
Yes, and it has already happened. As described earlier in this guide, a federal court ordered OpenAI in 2025 to preserve ChatGPT conversations that users had deleted, because they might be relevant evidence in unrelated litigation. Courts increasingly treat AI chat logs the same way they treat email and messaging records for discovery purposes.

What is the difference between local AI and edge AI.
The terms overlap heavily in casual use. Edge AI is the broader industry term for AI processing that happens on a local device, whether a phone, a sensor, or a desktop computer, rather than in a centralized cloud data center. Local AI, in the context of legal and business tools, typically refers specifically to running a capable language model on an office computer or small server, keeping documents and prompts entirely inside the organization’s own network.

Is local AI as accurate as cloud AI.
For open ended general knowledge questions, the largest cloud models currently hold an edge, though the gap is narrowing steadily. For tasks grounded in a firm’s own documents, contracts, financial records, or precedent files, a properly configured local model performs comparably for most practical purposes, because the model is working from material you provide rather than trying to recall the entire internet from memory.

How much does it cost to run AI locally compared to a cloud subscription.
Costs vary by the scale of the deployment and the hardware chosen, but the general shape is consistent: cloud AI is a recurring, usually per seat, monthly cost that scales upward as a team grows and as vendors periodically raise enterprise pricing at renewal. Local AI is primarily an upfront hardware and setup cost, after which ongoing costs are largely limited to electricity and occasional hardware refreshes, with the team retaining ownership of the hardware throughout.

What hardware do I need to run AI locally.
Requirements depend on the size of model and the workload, but a well specified modern desktop or small workstation with a capable graphics processor and adequate memory is sufficient for a wide range of legal and business use cases today, a bar that has dropped substantially over the past two years as smaller models have become more capable per unit of computing power.

Does GDPR apply to a small law firm or business using AI tools.
If your organization processes the personal data of anyone located in the European Union, including a client, a job applicant, or a website visitor, GDPR can apply regardless of your organization’s size or location. This is precisely why the destination and processing location of your AI vendor’s servers matters, and why a data processing agreement, not just a privacy policy, is the document that actually defines your exposure.

What should I do if an employee already pasted confidential data into a public AI tool.
Document exactly what was shared and when, review the specific vendor’s data deletion and training exclusion options for that account, and if the exposure involves client, financial, or personal data covered by a specific regulation, consult counsel about notification obligations, which vary meaningfully by jurisdiction and the type of data involved. Waiting to find out whether the exposure “matters” tends to be the costliest choice available, because most notification deadlines start running from the date of discovery, not the date the organization decides to act.

Is Microsoft Copilot, Google Gemini, or another major enterprise AI tool safer than ChatGPT.
Enterprise tiers of major providers generally offer stronger contractual protections than free consumer products, including training exclusions and defined data residency commitments, and are a meaningfully better choice than an unmanaged personal account. They still involve sending your data to a third party’s cloud infrastructure, which means every risk in this guide tied to vendor access, breach exposure, and legal discoverability still applies in some form, just with more contractual guardrails around it than a free tier offers.

Will local AI eventually replace cloud AI entirely.
Most industry analysts tracking this space expect a hybrid future rather than a full replacement, where routine, sensitive, or document grounded work runs locally, while occasional tasks that benefit from the largest frontier models still route to the cloud when the sensitivity of the material allows it. For legal and business teams handling privileged or financial data as a matter of course, the local portion of that hybrid setup is where the majority of day to day work increasingly belongs.

What is the single biggest mistake teams make with AI and confidential data.
Treating account setup as a one time decision rather than a policy. The most common failure pattern in every incident described in this guide is not a single dramatic breach, but an accumulation of individual employees making individually reasonable seeming choices, on personal accounts, without a policy telling them otherwise, until the pattern becomes a genuine liability nobody specifically decided to create.

Cloud AI and local AI, side by side

Sometimes the clearest way to see a decision is to lay the two paths next to each other rather than read about them in separate paragraphs.

QuestionTypical cloud AI accountTypical local AI deployment
Where does a prompt physically goTo a vendor’s data center, often across bordersStays on hardware inside your building
Can the vendor use your data for trainingOften yes by default, unless excluded by contractNot applicable, there is no vendor receiving the data
Who can access your conversationsVendor staff, contractors, and any subprocessor in the chainOnly people with access to your own network
What happens after you click deleteDepends on vendor policy, and can be overridden by a legal holdYou control deletion directly, on your own systems
Exposure if the vendor is breachedYour data is part of whatever the vendor lostNot applicable, your data was never on the vendor’s systems
Discoverable by subpoena to a third partyYes, as OpenAI’s litigation history showsOnly through a subpoena to your own organization, same as any internal record
Ongoing cost structureRecurring, usually per seat, subject to renewal pricing changesUpfront hardware and setup, modest ongoing cost, no seat based renewal
Works without an internet connectionNoYes, once set up
Who owns the infrastructureThe vendorYour organization

This is not a table designed to prove that one column is right for every situation. A team occasionally needing the single most capable frontier model available for a bounded, non sensitive research task may reasonably keep a cloud account for that narrow purpose. But for the daily, recurring work of a legal or business team, drafting from client files, reviewing contracts, analyzing internal financial data, the right column is the one that removes an entire category of risk from the table rather than managing it through a contract.

A glossary of terms worth knowing before you evaluate any AI tool

Hallucination. When a language model generates text that sounds fluent and confident but is factually wrong or entirely invented, such as a court citation that does not exist. Covered in detail earlier in this guide through the Mata v. Avianca case and the broader tracking database that followed it.

Shadow AI. AI tools used for work purposes without an employer’s knowledge, approval, or configuration, typically through a free personal account outside company oversight. IBM’s 2025 research found unsanctioned shadow AI use was a factor in a meaningful share of breaches studied.

Training data. The material a model learns from during its development, and in the case of many consumer AI products, potentially the material users type into the product itself, unless that use is contractually excluded.

Zero data retention. A specific type of contractual arrangement, typically available through business or API tiers, where the vendor commits to never storing a customer’s prompts or outputs at all, meaning the data cannot later be reviewed, retained under a legal hold, or exposed in a breach, because it was never kept in the first place.

Data processing agreement, often shortened to DPA. A binding legal contract, distinct from a general privacy policy, that specifies exactly how a vendor is permitted to process a customer’s data on their behalf, typically required under GDPR and similar regimes whenever personal data is involved.

Subprocessor. A third party a vendor relies on to help deliver its service, such as a cloud hosting provider or a customer support platform, which may also have access to or store a copy of customer data as part of that relationship.

Edge AI. The broader industry term for AI processing that happens on a local device, a phone, a sensor, a desktop computer, rather than in a centralized cloud data center. The global edge AI market, valued at roughly 24.9 billion dollars in 2025, is projected to more than quadruple by the early 2030s according to Grand View Research, discussed earlier in this guide.

Local AI, or on premise AI. In the context of this guide, running a capable AI model directly on hardware owned by the organization using it, keeping documents and prompts inside the organization’s own network rather than transmitting them to a cloud vendor.

Retrieval grounded AI, sometimes called RAG. An approach where a model answers questions using a defined, verifiable set of documents an organization provides, such as a firm’s own contract library or financial records, rather than relying purely on general knowledge from its training data. This approach meaningfully reduces hallucination risk for bounded, document based tasks, which is part of why it matters for the verification obligations described under Formal Opinion 512.

Attorney client privilege. The legal protection covering confidential communications between a lawyer and a client made for the purpose of seeking or providing legal advice. Sharing privileged material with an outside AI vendor, without appropriate safeguards, can put that protection at risk, distinct from and in addition to the separate ethical duty of confidentiality under Rule 1.6.

Where to start: a realistic path over the next ninety days

Fixing this does not require an overnight overhaul, and trying to do it all at once is usually why well intentioned policies never make it past the planning stage. A staged approach tends to actually get finished.

In the first thirty days, find out what is actually happening today. Most firms and business teams are surprised by the answer. Ask every team member, directly and without judgment, which AI tools they currently use for work, including personal accounts nobody formally approved. The goal of this step is not to assign blame, since as the Samsung case shows, this pattern shows up even among skilled, well intentioned professionals at sophisticated organizations. The goal is an honest inventory, because you cannot fix exposure you cannot see. Pair that inventory with the vendor checklist covered earlier in this guide for every tool currently in active use, and flag anything where the answers are unclear or concerning.

In the next thirty days, draw the line and communicate it clearly. Using the inventory from step one, decide which tools are approved for which categories of data, using the specific data categories named in the policy section above rather than a vague standard. Communicate this in plain language, in a short document, not a lengthy memo nobody will finish reading. Explain briefly why each restricted tool did not make the list, since a rule people understand is a rule people are far more likely to actually follow, and give people a clear, easy path to request review of a tool that did not make the initial list rather than quietly working around the restriction.

In the final thirty days, pilot a local alternative for your most sensitive recurring workflow. Pick the single task where sensitive data most often meets AI in your organization today, contract review, client intake summarization, financial report drafting, whatever it is for your team, and pilot a local AI setup specifically for that workflow rather than trying to replace every cloud tool across the organization simultaneously. A focused pilot lets your team build real confidence in the accuracy and usability of a local option before it becomes the default, and it gives you a concrete, low risk way to test the vendor checklist questions against a specific privacy first product rather than evaluating in the abstract.

By the end of ninety days, most teams following this path have gone from an unknown, unmanaged sprawl of personal AI accounts to a documented, defensible policy with a working local alternative already proving itself on real work. That is a meaningfully stronger position heading into 2027, when the regulatory and disciplinary consequences covered throughout this guide are, if anything, likely to keep tightening rather than relax.

A few more questions worth answering

Does using AI locally mean giving up newer model improvements entirely.
No. The open model ecosystem that powers local AI tools is updated frequently, and a well built local product lets your team pull in a newer model on a schedule your firm controls, after a short internal review, rather than having a cloud vendor swap the model behind your account without notice, which has happened often enough across major providers that many enterprise users have learned to expect it.

Can a local AI setup still integrate with the practice management or accounting software my team already uses.
Generally yes, though the specific integrations available depend on the product. Purpose built local AI tools designed for legal and business use are typically built with the expectation that they need to fit into an existing workflow, document management systems, practice management platforms, accounting software, rather than replace them, since the goal is removing a specific risk, not disrupting a team’s existing operations.

What happens to accuracy if the model never sees the broader internet the way a cloud model does.
For general open ended questions, a cloud model’s broader training exposure can be an advantage. For the work most legal and business teams actually need day to day, reviewing a specific contract, summarizing a specific deposition, answering a question about a specific company’s own financial records, the model performs best when it is grounded in the actual document in front of it rather than trying to recall something similar from general training data, which is exactly the kind of task a well configured local, document grounded setup is suited for.

Is it realistic for a solo practitioner or a five person team to run this, not just a mid size firm.
Yes, and this is precisely the audience purpose built privacy first local AI products are increasingly designed for. The barrier that used to make this exclusive to large organizations with dedicated IT staff, expensive specialized hardware and a steep technical setup, has fallen substantially over the past two years, and a small team’s exposure per client relationship is, if anything, higher than a large firm’s, given fewer safeguards to fall back on if something goes wrong, which makes this a case where a small team arguably has more to gain from making the switch, not less.

The numbers from this guide, in one place

For anyone who wants the headline figures without re reading every section, here is the full set gathered together, each one sourced above.

  • 4.44 million dollars, the global average cost of a data breach in 2025, and 10.22 million dollars, the US average, according to IBM’s Cost of a Data Breach Report.
  • 97 percent, the share of AI related breaches IBM studied where the organization had no access controls in place on the compromised AI system.
  • 670,000 dollars, the average additional cost IBM found among breaches involving heavy unsanctioned shadow AI use, on top of an already elevated baseline.
  • 77 percent, the share of enterprise employees using generative AI who had copied and pasted data directly into a prompt, per LayerX’s 2025 research, with 82 percent of those pastes coming from personal accounts outside employer visibility.
  • 34.8 percent, the share of corporate data pasted into AI tools that Cyberhaven classified as sensitive in 2025, up from 10.7 percent just two years earlier.
  • 38 percent, the share of US workers who told Kolmogorov Law’s 2026 survey they had entered work information into a personal AI account their employer did not control.
  • Over 1,600, the number of court cases worldwide documented by Damien Charlotin’s tracker as of June 2026 involving AI hallucinated material submitted in legal filings, up from roughly 200 a year earlier.
  • 15 million euros, Italy’s Garante fine against OpenAI in December 2024, the first GDPR penalty issued against a generative AI company.
  • 29 percent, the share of law firms that told the ABA’s 2025 TechReport survey they had experienced a security breach at some point.
  • Four times, roughly how much more often small and mid size businesses experience confirmed breaches compared with large organizations, per Verizon’s 2025 Data Breach Investigations Report.
  • 118.7 billion dollars, the projected size of the global edge AI market by 2033, up from 24.9 billion dollars in 2025, according to Grand View Research, a growth curve driven in part by rising demand for data privacy and reduced dependence on centralized cloud infrastructure.

Read individually, each number describes a single incident, a single survey, a single market forecast. Read together, they describe a single trend line pointed in the same direction: the cost of getting this wrong keeps climbing, the number of organizations getting it wrong keeps climbing alongside it, and the tools built to sidestep the problem entirely are growing fastest of all.

The bottom line

Every risk covered in this guide traces back to the same root fact: the moment sensitive information leaves your building and lands on infrastructure someone else owns, you have handed over a decision that used to be entirely yours. That does not make cloud AI reckless to use for every purpose. It means the sensitivity of what you are handing over should determine where it gets processed, not convenience alone, and for a large share of what legal and business teams actually do with AI day to day, drafting from client files, reviewing contracts, working with financial records, that sensitivity is high enough that the calculation should tilt toward keeping the work in house.

The pattern across every case study in this guide, Samsung’s engineers, the Mata v. Avianca sanctions, Italy’s fifteen million euro fine, the federal court order to preserve deleted ChatGPT conversations, is the same pattern repeating with different names attached. Skilled, well intentioned professionals used a convenient tool without fully understanding where their information was actually going, and found out the hard way, months or years later, that convenience and control were never the same thing. None of them set out to create a liability. They simply never asked the second question, where does this go, before they asked the first, will this save me time.

Heading into 2027, the regulatory floor, the disciplinary consequences, and the sheer volume of documented incidents all point in one direction: the organizations treating this as a genuine architecture decision, not a settings toggle, are the ones that will spend the next few years building on a foundation instead of managing an accumulating liability. Running AI locally, on hardware your firm or your business actually owns, is not the only way to get there. It is, for the material covered in this guide, privileged client files, financial records, competitively sensitive strategy, the most direct one, because it removes the question of trust entirely rather than asking you to manage it.

If your team handles the kind of data covered in this guide and you have been putting off a real evaluation of where your AI tools actually send it, this is a good moment to stop putting it off. A local, privacy first AI assistant built specifically for legal and business teams, one that keeps every prompt, every document, and every client detail on hardware you own and never transmits any of it to a third party, is exactly the kind of foundation worth building on before your next cloud AI renewal locks your team in for another year. (Add your product name, a short description, and a link here before publishing.)

Sources and further reading

Every claim and statistic in this guide is drawn from a named, dated, publicly available source, listed below for anyone who wants to verify a figure or read further into a specific incident or regulation.

Leave a Comment