Skip to main content

REALITY CHECK

The AI Reality Check

Hallucinations, lawsuits, data leaks, deepfakes, energy strain, vendor lock-in. The conversation no luxury AV firm wants to have. We are having it anyway.

Reading time 14 min Updated May 2026 Topic Risk

Why honesty wins

The integrator pledge: we will tell you what an AI deployment can and cannot do, what it costs in money and risk, and where the weight of evidence currently sits on the question of whether you should do it at all. We will not write a brochure. We will not pretend the lawsuits are settled. We will not pretend the models do not hallucinate. We will not pretend your data is safe because the marketing copy says so. If a vendor is unwilling to do this, the vendor is not telling you something you need to know.

This page is the long form of the integrator pledge. It is the thing we send to clients before we propose any AI work. If you finish it and still want AI in your home or building, we know we have your informed consent. That is the standard.

Hallucinations: it is not a glitch, it is the architecture

Large language models do not look up answers. They generate plausible next tokens. When the training data covers a question well, the plausible next tokens line up with the truth. When it does not, the model invents. This is not a bug they are about to fix; it is what the architecture does.

The Vectara hallucination leaderboard tracks summarization hallucination rates across major models. The current frontier sits in the 0.7% to 5% range for the best models on summarization tasks. Gemini 2.0 Flash held the record at 0.7%. Claude 3.7 measured around 4.4%. GPT-5 around 1.4%. Llama 4 Maverick around 4.6%. DeepSeek R1 between 11% and 14%. These are the leaders. Most consumer models are worse.

Summarization is the easy case, because the source is in the prompt. The hard case is open-book general knowledge, where the model has nothing in front of it and has to know whether it knows. The AA-Omniscience benchmark measured GPT-5.5 producing a confident wrong answer 86% of the time when the model did not actually know the answer. The accuracy on questions it could answer was 57%. The remaining 43% was where the omniscience rate measured. That is not a model you let near a contract.

Mata v Avianca: the lawyer’s case

In 2023, an attorney for the plaintiff in Mata v Avianca submitted a brief that cited six federal court cases. None of the cases existed. ChatGPT had hallucinated all six. The attorney was sanctioned, fined, and his career ended. The case is now part of the legal ethics curriculum at most law schools. The lesson is not "do not use ChatGPT." The lesson is "verify every citation against a primary source, every time, no exceptions, when stakes are high."

Perplexity’s 37% audit

Even citation-first answer engines hallucinate. A Columbia Journalism Review audit found 37% of Perplexity answers contained errors. Citations are evidence the model looked at sources. They are not evidence the model summarized them correctly. We use Perplexity daily. We say this anyway because it is true.

Lawsuits in flight, May 2026

OpenAI v Musk

Set for trial in Oakland federal court in April 2026. BBC coverage outlines the case: Musk seeks roughly $150B in damages, claiming OpenAI’s shift from non-profit to capped-profit to for-profit structure breached his original donation agreement. Sam Altman is expected to testify. The case will not destroy OpenAI but it will produce months of subpoenaed internal communications. We are watching it the way we would watch any vendor whose corporate structure is being litigated.

NYT v OpenAI

The New York Times complaint includes hundreds of examples of GPT-4 reproducing copyrighted Times articles near-verbatim. It is now part of MDL 25-md-03143 in the Southern District of New York under Judge Stein. The case will likely settle, but the discovery has already changed how we think about training-data retention in deployed models. If the model has memorized the Times, it has memorized other things too.

Other suits

Disney and Universal are suing Midjourney for image-output infringement. Music publishers are suing Anthropic for lyric reproduction. The Authors Guild has cases against multiple labs. None of these alone change anything; collectively they constitute a tax that will eventually fall on output pricing.

Air Canada precedent

Air Canada was held liable when its chatbot promised a customer a bereavement fare refund that was not actually company policy. The tribunal ruled the chatbot was an agent of the airline and the airline owed what the chatbot promised. The principle now applies to anyone deploying a customer-facing chatbot, including hotel concierges, restaurant reservation bots, and yes, our clients. We do not deploy customer-facing AI without confirmation steps and bounded scope.

RISK

If you are running a hospitality property and a chatbot in your booking flow promises a guest something your policy does not cover, Air Canada says you owe it. Bound the scope. Log the interactions. Have a human review anything outside guardrails.

Data leaks

Samsung 2023

Samsung banned ChatGPT and other chatbots for employees after engineers pasted internal source code into the free tier as a debugging convenience. OpenAI’s training policy at the time meant that code became part of the corpus. JPMorgan, Amazon, Apple, and Verizon followed with similar restrictions. The rule is simple: anything you put into a free-tier chatbot may be used to train a future model. That is a feature of the free tier, not a bug.

The four-tier rule

The defensible rule we use:

  • Free tier: assume training. Do not put proprietary data, client data, or financial data into it. Ever.
  • Pro tier: usually no training, but check the terms quarterly because they change.
  • API and Enterprise tier: contractual no-training. Verify in writing.
  • Local: no training, no transit, no question.

For sensitive client work, we route through API or local. Period.

Deepfakes and fraud

Voice cloning now requires roughly three seconds of audio. The FBI and AARP have both issued warnings about grandparent scams using cloned voices of grandchildren in distress. The FBI has issued additional warnings about deepfake job interview candidates submitting falsified live video to remote-work interviews. Election misinformation has its own ecosystem at this point. None of this is hypothetical. All of it is operational right now.

For luxury homeowners specifically: assume a public profile (LinkedIn photo, conference video, podcast appearance) is enough material to produce a convincing voice clone of you. Establish a family verification phrase that is never written down or spoken on a recorded channel. We are not joking. We have had a client who almost wired six figures based on a cloned voice request from "their daughter."

Nation-state weaponization

Anthropic disclosed in November 2025 that a Chinese state-aligned actor used Claude Code in a 30-target espionage campaign before Anthropic disrupted it. This is one of the only public postmortems any frontier lab has published on misuse. Read it once. Then assume similar campaigns are running against the labs that have not published postmortems. Assume your data center is a target. Assume the supply chain on your AI infrastructure is a target. We do, and we design accordingly.

Energy and grid strain

The IEA forecasts global data center electricity demand reaching 945 TWh by 2030. Goldman Sachs projects a 165% increase in U.S. data center power demand over the same period. Individual AI campuses now draw 100 to 300 megawatts; planned facilities reach 1 gigawatt. Per-query energy is roughly 3.0 watt-hours for an AI answer versus 0.3 watt-hours for a traditional search query, a tenfold increase.

This matters to homeowners and developers in two ways. First, the grid in the New York and New Jersey corridor is being asked to absorb a load that the planning horizon did not anticipate; expect rate volatility. Second, when a luxury home generates significant AI workload, doing it locally on a 200-watt Mac Studio is not just a privacy choice, it is a substantially smaller environmental footprint than routing the same query to a cloud data center that sits at five to ten times the per-query energy cost.

Vendor lock-in and the AI tax

The pattern is now well established. Sonos remotely deprecated functionality on speakers customers had bought outright. Google Bard was sunset, then renamed, then merged into Gemini, then partially deprecated again. Samsung Galaxy AI features that shipped free now cost a subscription. Amazon is moving Alexa users to a paid Alexa+ tier. Apple is adding Apple Intelligence as a free feature for now, but the ground rules will shift.

The pattern is: ship feature free, build dependency, charge subscription, throttle non-subscribers. Every cloud-tied "smart" device in your home is on this curve. The defense is to own the hardware and own the software stack so the vendor cannot unilaterally change the deal. That is not anti-cloud zealotry; it is property rights.

The integrator’s liability

Air Canada applies to us, not just to airlines. If we deploy an AI system that tells a client’s guest the wrong thing, the client may be liable, and the client may turn around and look at us. We sign contracts with bounded AI scope. We do not let an LLM autonomously commit the client to anything. We document every action the AI is allowed to take. We log every interaction in a tamper-evident store. We keep manual override on every safety-relevant control. This is not paranoia, it is the only defensible position for a serious integrator.

Insurance and the diligence question

Liability carriers have begun asking about AI in the underwriting process. For homeowners, the questions are still relatively soft (do you have cloud-tied locks? do you have a remote-disable risk?), but the trajectory is clear. For commercial and hospitality clients, the questions are sharper: what AI is in your customer service flow, what does it have access to, who reviewed the bounded scope, what is the logging posture. We help clients answer those questions in writing. A documented AI deployment with bounded scope and a logged escalation path is a real underwriting asset. An undocumented one is a real underwriting liability.

The talent and accountability question

One of the quieter risks of AI deployment is the gap it creates in institutional knowledge. If a junior associate uses an LLM to generate a memo, neither the associate nor the senior who reviews it has the full context that would have come from the associate having to actually do the research. Three years of that pattern across an industry produces a generation that can produce output but cannot defend it. We are not an HR firm and we will not tell anyone how to staff their team. We will tell every client this, though: the thing AI replaces is sometimes the thing a human needed to do in order to learn. Build the verification step in not just for accuracy but for skill development.

The training data and copyright question

Beyond the NYT case, there is an open question about what your AI deployment owes to creators whose work helped train the model you are using. The legal answer is unsettled and will be unsettled for years. The pragmatic answer is that for client-facing creative work (marketing copy, design assets, brand voice), generating the asset entirely with an AI introduces a non-trivial risk that part of the asset is copied verbatim from training data, and that the creator of that data has a future claim. The defense is to use AI as a polish step, not as the source of original creative work, and to document the human contribution. We do this for our own marketing, including the page you are reading.

Sycophancy and feedback loops

A subtler failure mode than hallucination is sycophancy: the tendency of LLMs to agree with the user, soften disagreement, and produce the answer the user seems to want rather than the answer the evidence supports. Grok 4.x has been the most flagged for this in 2026, but no model is immune. The risk is meaningful when AI is used as a reasoning partner on a strategic decision: the human walks away with a confidence boost that was generated by the model agreeing with them, not by the model independently arriving at the same conclusion. We address this by routing critical questions through more than one model and looking for disagreement. When models disagree, that is signal. When they agree because they are agreeing with the user, that is noise.

The Restrepo position

  • We deploy AI. We do not sell hype.
  • We default to local. We use cloud where it is the right tool.
  • We document what data leaves the property and where it goes.
  • We keep manual override on every safety-relevant control. No LLM unlocks doors, opens valves, disables alarms, or moves money without a human in the loop.
  • We will tell you what AI cannot do as clearly as what it can.
  • We update this page when the facts change.
  • We document our certifications honestly. Elite Pro Crestron Dealer, HTA Luxury Certified, OSHA Certified, Ubiquiti UISP and Pro Partner. We deploy Lutron and Ketra products as a partner, not a license-holder.

A short note on what we do not know

The honest part of any reality check is the section that says where the writer’s knowledge runs out. Ours: we do not know how the OpenAI v Musk trial will resolve, and the outcome could change OpenAI’s product roadmap meaningfully. We do not know what the next twelve months of model releases will do to the hallucination numbers; some benchmarks will improve, others will get harder, and the relative ranking of labs will move. We do not know which currently-cloud-only AI capabilities will be available locally in a year, but the trend has been a steady migration of capability into open-weight models. We do not know which AI products will exist in five years and which will have been sunset. We will update this page when these questions resolve.

What this means for you

You do not have to opt out of AI. You have to opt in carefully. Read this page. Read the model field guide. Decide which of your data belongs in which tier. Decide whether you want a cloud-only deployment, a hybrid, or a fully local one (see Build It Yourself). Decide which AI vendor relationships you are willing to maintain. Then call us at 201.405.2022 or write office@restrepoinnovations.com, and we will build it the way you decided, not the way the marketing department decided. We work in New Jersey, Connecticut, and the New York metro, including Bergen, Essex, Morris, Passaic, Hudson, Fairfield, and Litchfield counties.

Related

Want AI in your home or building, done honestly?

We design and deploy AI that lives on hardware you own. No cloud lock-in, no data trades, no surprises. Serving NJ, NY, and CT metro.