Skip to content
Back to Journal
Engineering3 min read26 January 2023

GPT-3 to GPT-4 — how we actually used AI in client work (not the hype version)

GPT-3 to GPT-4 — how we actually used AI in client work (not the hype version)

When GPT-3 came out, we did what everyone did. We played with it, marveled at it, and then tried to figure out if it was actually useful. The honest answer in 2021 was: sometimes, for certain things, with a lot of hand-holding. We used it to generate first-draft ad copy variations for A/B testing. We used it to help write product descriptions for an e-commerce client who had 400 SKUs and no time to write them properly. Those use cases worked.

What did not work: asking it to write anything that required brand voice consistency, anything technical about a specific client's business, anything that needed real accuracy. The hallucinations were a real problem. We spent more time fact-checking than writing in some cases.

GPT-4 changed the calculus significantly.

What actually changed with GPT-4

The reasoning quality went up enough to matter. We started using it for a few things that had not been practical before. First, drafting structured content outlines for clients who needed blog content but had no internal writers. The outline quality was high enough that a junior team member could execute the actual writing without constant supervision. Second, generating schema markup JSON-LD for client pages, which is tedious and error-prone when done manually. GPT-4 got it right most of the time and made a class of errors that were easy to spot.

Third, and this surprised us: using it as a first-pass QA tool. We started pasting landing page copy into a prompt and asking it to flag claims that seemed vague, unsubstantiated, or unclear to a first-time reader. This did not replace human review but it consistently caught things we had read past.

What we stopped pretending AI could do

We had a client who wanted us to use AI to generate their monthly SEO blog content end-to-end. No human writer involved. They had seen the demos and thought it was solved. We tried it for two months. The content was technically coherent. It ranked for nothing. The reason was not a mystery. It had no original perspective, no data from their actual business, no examples from their industry experience. It was plausible text that said nothing.

The clients who got real value from AI in their work were the ones who treated it as acceleration, not replacement. The lawyer who used it to draft the first pass of routine documents and then edited them. The product company that used it to summarize user feedback and spot patterns. The marketing team that used it to generate ten headline variants and then chose the two worth testing.

We still use it every week. We are honest with clients about where it helps and where it does not. That is the unglamorous version of the AI story, and it is the one that actually holds up.

Published 26 January 2023
Start a Project