Andrej Karpathy has done it again. Not necessarily invented something new. Something arguably more valuable. He noticed a behavior that many people working intensively with AI already do, consciously or otherwise, gave it a name, and therefore ensured that we’ll all be discussing it for the next six months.
The term is rambling. The idea is wonderfully unsophisticated: Instead of spending ten minutes constructing the perfect prompt, turn on voice input and spend ten minutes talking absolute rubbish.
Well, structured rubbish. Actually, don’t even bother with the structured part.
Your Prompt Is Probably Missing Most of the Information
Karpathy’s observation is that LLM performance often isn’t limited by how elegantly you formulate the request. It’s limited by how many useful bits about your actual intention reach the model.
Suppose you type: Write a landing page for this product. Perhaps 50 words follow. But inside your head there are another 5,000.
You know why you’re building it. You know who the customer is. You remember what you tried last month. You know which competitor’s positioning annoys you. You know the sentence you absolutely don’t want. You know that management technically asked for X but actually needs Y. You know what “good” looks like.
The model knows none of this. So it does the only reasonable thing. It guesses. And because LLMs have become extremely good at guessing, the first result often looks surprisingly competent.
Unfortunately, it’s competent at solving the problem it inferred, which may not be the problem you actually have.
Then begins the familiar ritual:
No, less corporate. That’s not quite what I meant. Keep the previous structure. Actually, forget that. The audience isn’t developers. Don’t emphasize the AI part. This is closer, but…
Twenty messages later you’ve finally transmitted the context that could have been provided at the beginning. Congratulations. You have invented a very expensive compression algorithm.
So Don’t Compress
Rambling reverses the process. Lean back. Turn on speech recognition. Talk. For five minutes. Ten minutes. Jump between ideas. Repeat yourself. Change your mind halfway through a sentence. Mention something apparently irrelevant and then explain why it isn’t.
Say:
“Basically… well, actually, that’s not quite right…
the real problem is…”
This is precisely the stuff we usually delete while constructing a “good prompt.” And that may be a mistake. Because those fragments contain information.
The hesitation tells the model where the uncertainty is. The correction reveals what distinction matters. The tangent supplies context. The example reveals your implicit quality threshold. The complaint tells it what to avoid. The contradiction may expose the actual problem.
You’re not writing instructions anymore. You’re performing a context dump.
The Internal Editor Is the Enemy
There’s another subtle point here. You can technically ramble in writing. I do. But voice changes the bandwidth.
When typing, we automatically edit. We shorten. Organize. Remove repetitions. Turn fuzzy thoughts into declarative sentences. Essentially, we compress ourselves before the model ever sees the source material. Speech makes this harder. You speak faster than you type, and the friction of editing disappears. That means more of the intermediate reasoning survives transmission.
The objective isn’t beautiful language. The objective is lossless-ish transmission of intent. Let the LLM perform the compression afterward. It is, after all, rather good at that.
Prompt Engineering Is Becoming Context Engineering
This also fits a broader pattern. Early LLM usage revolved around prompts:
You are an expert copywriter…
Then came increasingly elaborate system prompts. Then RAG. Memory. Tools. Agent harnesses. Persistent project context. Skills. Personalization. And now the important question increasingly isn’t:
What magical sentence should I type?
It’s:
What does the model need to know in order to understand what I’m actually trying to accomplish?
That’s a much healthier question. And rambling is simply the lowest-tech implementation of context engineering imaginable.
No vector database. No knowledge graph. No orchestration framework. Just a microphone and your inability to shut up.
Beautiful.
I’ve Been Doing This More and More
I’ve noticed the same thing in my own work. Especially with coding agents.
For a complicated task, I increasingly don’t want to spend time producing a beautiful specification before the agent has seen the problem.
I’ll explain what I’m thinking. What’s broken. What worries me. What I’ve already tried. What should definitely not change. What might be related. What I suspect is irrelevant but perhaps isn’t.
Then I let the model reconstruct the problem. Often its cleaned-up interpretation of my rambling is clearer than the version I had in my head. And then we can work from that.
It’s almost like using the model as a cognitive compiler.
Input: human thought spaghetti
Output: reasonably typed specification
Karpathy’s Actual Superpower
I’m slightly jealous of Karpathy’s ability to do this. Not the rambling. I’m perfectly capable of that.
The ability to observe something everyone has begun doing informally and say: This is a thing. It needs a name.
Vibe coding followed a similar trajectory. The behavior existed.
Someone articulated it. The articulation created the category. The category changed how people thought about the behavior.
Now perhaps we’ll get rambling. Soon there will presumably be rambling frameworks. Rambling benchmarks. Enterprise Rambling Platforms. Someone will raise $38 million for a rambling observability layer. Gartner will produce a Magic Quadrant. We have been warned.
There Is, However, Another Explanation
Put on the aluminium hat for a moment. There is an alternative interpretation of the entire evolution of AI UX.
Vibe coding: Generate more code.
Agents: Let the model run for hours.
Extended thinking: Let the model think longer.
Rambling: Please provide ten minutes of additional tokens before we even start.
An extraordinary coincidence. Almost every major conceptual breakthrough in AI interaction appears to have one curious side effect: the token meter spins faster.
I’m not saying Big Token is secretly manipulating Andrej Karpathy into convincing us to dictate our stream of consciousness directly into inference clusters. Obviously. That would be ridiculous. Probably. But if next month’s hot AI trend is called “continuous ambient context streaming,” I’m buying the foil in bulk.
The Useful Part
Conspiracy theories aside, I think rambling describes something genuinely important. The bottleneck in working with increasingly capable models is shifting. The model often has enough intelligence to perform the task.
What it lacks is your state. Your assumptions. Your history. Your preferences. Your constraints. Your definition of success.
The expensive part isn’t always getting the model to reason harder. Sometimes it’s simply transferring enough of yourself into the context window that it stops having to guess. So next time you’re staring at an empty prompt box trying to formulate the perfect instruction, don’t.
Press the microphone button. Talk for ten minutes. Make a mess. Then ask the machine to figure out what you meant. Apparently, that’s prompt engineering now.

