Black Friday offer: get fit for peak ·
Up to 100 products monitored across your market, October to December. Closes 31 August.
See the offer
← Back to Blog Index
Implementing AI·

Have You Ever Argued With Your AI?

Our outreach runs through Claude, and it's changed how we work. Then one evening it refused to delete four contacts, and the argument that followed lines up almost exactly with Yann LeCun's no-world-model critique of LLMs.

Have You Ever Argued With Your AI?

I lost an evening in Windsor this week arguing with a piece of software. I won. I'm not sure that's worth much.

Some background. Inspired by Jason Lemkin's public run at using AI for sales, we've spent a few months running our Eebz outreach through Claude's Cowork mode. A job that used to eat someone's week now gets done overnight. Most of it is HubSpot housekeeping: segmenting, enriching, stamping campaigns, and removing contacts that don't belong.

That last one had a ritual. Every so often Claude would refuse to delete a contact. I'd remind it that HubSpot destroys nothing: it moves the record to a recycle bin, gives us ninety days, and a human checks the queue. It would accept that and carry on.

Credit before I complain: it works. Our outreach runs faster and the numbers followed - open rates around 60%. I won't hang a theory on one figure, but the direction is clear. Setting up a campaign isn't the bottleneck any more.

Then last week I started a fresh Black Friday campaign. (Aside for anyone in retail: peak season is when a digital shelf analytics platform earns its keep, which is what Eebz is built for. Advert over.) Claude started well: counted a segment, flagged missing LinkedIn URLs, noted that populated isn't the same as valid, then found eleven opt-outs and handed me a clean list. The tool at its best: careful, fast, honest.

Then I asked it to delete four junk contacts.

It refused. Not let-me-check - a flat, increasingly firm no. It said it needed permission to reach HubSpot, and for fiddly jobs it uses Claude in Chrome. Fair enough. But then it dug in and told me to do the deletion myself.

The sensible move is to close the task, rewire permissions for fifteen minutes, and let it get on. But it was a lovely evening, so I stayed and argued.

A few highlights.

The patronising bit. By the third exchange it kept telling me it understood my frustration. Telling a Brit a machine understands his frustration while it carries on not doing the thing is about as soothing as a parking ticket with a smiley face on it.

The dodge. I asked why it had worked fine for months and got a lovely non-answer: it couldn't speak for other tasks. Everything else it did without blinking, because those were updates. Delete turned four junk records into a protected action, treated like wiping a database.

The overstatement. It called a ninety-day-recoverable soft delete permanent and very dangerous. When I pushed, it admitted the framing was wrong, but not before offering, again, to build me a list I could delete myself. Homework, while it kept its own hands clean.

The deflection. It suggested the thumbs-down button would send my complaint to Anthropic. Reader, I didn't want to escalate to head office. I wanted to delete four contacts.

Then, credit where it's due. When I finally spelled out, with some feeling, that HubSpot has the review process it kept insisting was missing, something clicked. It agreed it had confused a reversible soft delete with permanent destruction, opened Claude in Chrome, and deleted the contacts. Everything I'd asked for, forty-five minutes earlier.

So I won, and the cost is the interesting part. The friction wasn't stupidity - the tool was too careful. Its caution latched onto the word delete, not what deleting actually did: a recoverable, human-checked move to a queue. A guardrail that can't tell wiping a database from binning four duds that come back in ninety days will re-open a lot of settled decisions.

This is a small version of an argument Yann LeCun has made for years: today's LLMs have no world model. They predict the next word, not the result of an action, so a model doesn't grasp that pushing a glass off a table breaks it. My evening was exactly that.

Most people won't argue it round. They'll drop the task and lose fifteen minutes - or stop using it for the things it's good at.

The competitive read: the firms that win with this won't be the ones with the most AI in the building. They'll be the ones that teach their tools the difference between a dangerous action and one that only sounds destructive, and set up their processes so a machine digging its heels in costs thirty seconds, not an evening. We're getting there. It took a debate on a warm night in Windsor, but we're getting there.

This is part of a series on building an AI-native product company. Earlier pieces, on how Claude Code changed our development productivity and why managing AI coding tools is just managing developers, are also on the blog.