Writing

It Was Never Just the Language Model

You're not paying for an LLM. You're paying for the system that makes it useful.

Copy link to “Text Is the Model’s Oxygen”Text Is the Model’s Oxygen

On its own, a language model does just one thing. It reads some text and guesses what word should come next. Then it adds that word and guesses again, one small piece at a time, until it has a full answer. That’s the whole job.

It can’t open a file or run a program. It doesn’t remember what you told it yesterday or know today’s date. It can’t even check if its own answer is right. And no, it can’t tell you why your ex cheated on you, but it will happily guess.

Text goes in, text comes out, and without text it can do nothing at all. That’s why text is the model’s oxygen.

But to be truly useful, a model needs more than oxygen. A lot more.

Copy link to “So What Is a Harness?”So What Is a Harness?

Think about a horse. It’s strong. It can run fast and pull heavy things. But leave it alone in a field, and what does it do? It eats some grass, runs around a bit, and takes a nap. All that power, and none of it is helping you.

Now put a harness on it. A harness is the set of straps and gear that connects the horse to a cart or a plow. It lets you steer the horse and point all that power in one direction. Suddenly the same horse is carrying goods to town or plowing a whole farm. The horse didn’t get stronger. It just got a harness.

But the harness has to fit well. A bad one makes the horse pull the wrong way, or even hurts it. And without a horse, a harness is just a pile of straps. The horse brings the power. The harness brings the direction.

That’s exactly where the AI world got the name. A language model is the horse: smart and powerful, but on its own, it just turns text into more text. The harness is all the software built around it. It decides what text goes in, what happens with the text that comes out, and how to keep the model on track until the job is done.

So when an AI tool feels smart, it’s not just the model. It’s the model and its harness working together.

Copy link to “Harness Plays Middleman”Harness Plays Middleman

The harness sits between you and the model, and it’s a big reason the model feels almost human to talk to.

Think of it as a translator working both ways. When you ask for something, the harness packs your request together with everything else the model needs to know. When the model replies, the harness unpacks that reply and acts on it, whether that means showing you an answer or getting something done.

Doing this well takes a lot of moving parts. Getting them right is a skill of its own, called harness engineering. Here are the most important parts of a harness, apart from the model itself.

Copy link to “Context”Context

Context is everything the model can see at one time: your message, the chat so far, any instructions, and any outside information the harness adds. The model doesn’t remember your conversation. Each time you send a message, the harness sends the whole chat again, and the model reads it fresh. There’s also a limit to how much text fits at once, called the context window. Fill it with clutter and the model gets worse, so the harness has to choose what goes in and what stays out.

Copy link to “Tools”Tools

Tools are things the harness can do on the model’s behalf, like searching the web, reading a file, or running code. The harness shows the model a list of these tools, each with a name and a short description. It’s like the apps on your phone: each one does a single job, and the model can see what’s available and what each one is for.

Copy link to “Tool Calling”Tool Calling

When the model wants to use a tool, it doesn’t do anything itself. It just writes a request, like “search the web for today’s weather in Paris.” The harness runs the real search and pastes the results back in as text, and the model carries on. It asks, and the harness does the work.

Copy link to “Loop”Loop

Real tasks take many steps, so the harness runs a loop: ask the model, run any tool it requested, send back the result, and repeat until the model says it’s done. The catch is that the model decides when it’s done, and sometimes it’s wrong. That’s why good harnesses also set limits, so the loop can’t run forever.

Copy link to “Memory”Memory

Since the model forgets everything, the harness remembers for it. During a task, it keeps track of what’s been done and what’s left, which is often called the state. Between sessions, it saves notes, to-do lists, or progress files to read next time, which is what we usually call memory. It’s a bit like leaving sticky notes for your future self.

Copy link to “Guardrails”Guardrails

A model that can take actions can also take the wrong ones, like deleting the wrong folder. Guardrails are the rules that prevent this. The harness can ask you before risky actions, block some actions entirely, or run everything in a safe, closed-off space. It also has to be careful with text from the web, which can contain sneaky instructions meant to trick the model.

Copy link to “The Model Is Lonely”The Model Is Lonely

The model is lonely. Kind of like you, dear reader. Just kidding.

But it really is. On its own, the model sits in a dark room with no windows. It can’t see what’s happening in the world, it can’t touch anything, and it can’t check if it’s right. The harness gives it all three: eyes to see the world, hands to take action, and a way to find out whether it got things right.

Think of Ben 10. Ben is an ordinary kid with no superpowers, until he finds the Omnitrix and suddenly can do amazing things. The model is Ben, and the harness is its Omnitrix. But the Omnitrix only works its magic when it fits well. A badly built harness holds back even the smartest model, and a well-built one helps it do wonders.

So the next time an AI tool impresses you, remember: it was never just the language model.