Aurora AITell us your case

Offering

ServicesProductsCase studies

For whom

Private EquityEnterpriseSMB
ServicesProductsCase studiesAboutBlogContact

Knowledge base

Start hereWikiGlossaryGuides

Working with Claude Baza wiedzy

How to brief an AI agent to work unsupervised — and know when it's done

An autonomous agent is only as good as the brief you hand it. Here's the anatomy of a brief an agent can run unsupervised, and how it knows when it's done.

A single sheet of instructions unfolds into an ordered network of glowing nodes running on its own; dark graphite background, cool blue shading into green.
Pojedyncza kartka-instrukcja rozwija się w uporządkowaną, samodzielnie płynącą sieć świetlnych węzłów; grafitowe tło, chłodny błękit przechodzący w zieleń.
Working with Claude#ai-agents #agentic-workflows #prompting #autonomous-agents #definition-of-done

You can hand a capable AI agent one big, ambiguous goal — "build me a complete small business from scratch, starting with nothing but the open internet" — then walk away and come back to a finished package: market research, a chosen problem, a product, a brand, a landing page, even a launch video, all tied together in a single recap you can open and follow. That's not hypothetical — a single unattended run really can carry an empty folder to something you could take to market. The reflex is to credit the model for it, or to go hunting for the one clever prompt that supposedly unlocked it. Both miss what actually happened. What made walking away safe wasn't the model and wasn't a magic sentence. It was the brief.

So I'll take the brief apart. First, what an autonomous brief actually contains, part by part, with the one part that quietly does more work than all the others. Then how you tell an agent to organise its own work, so the result comes back stronger than a single pass could manage. And at the end, an honest word on what this doesn't do.

The anatomy of a brief an agent can run without you

A brief you can leave running has parts, and each one has a job. Skip a part and you either can't walk away or you come back to the wrong thing.

The mission is the goal stated as an outcome, not as a set of steps. "Find a real, painful, underserved problem that people are complaining about right now, design a business around it, and hand me a finished package I could take to market — then prove to me why it would work." Notice what that isn't: it's not a task list. You describe the result you want and let the agent find its own path there. And you give it honest room to be good, not just safe — a line like "within the limits below you have total creative freedom; I want your best work, not your safest work" is not decoration. An agent kept on a short leash produces cautious, forgettable output. Told it has real latitude, it uses it.

The guardrails are the hard limits that make leaving the room safe. Three carry most of the weight. Spend nothing new — the agent works only with what it already has. Publish nothing — everything is built locally, so a mistake stays on your machine instead of going out into the world under your name. Invent nothing — every fact and figure has to be checked against a real source, so the agent can't paper over a gap with something plausible it made up. Add one more of your own: give it only the keys it needs. Whatever's already wired into the project is fair game; nothing beyond that. Guardrails aren't the boring part of the brief. They're the part that lets you leave.

Abstract cross-section of a document split into layers: the goal on top, boundaries at the sides, an ordered run of stages in the middle, a clear sign-off line below; graphite, blue and green.
Abstrakcyjny przekrój dokumentu rozłożonego na warstwy: cel u góry, ramy po bokach, uporządkowany ciąg etapów w środku, wyraźna linia odbioru u dołu; grafit, błękit i zieleń.

The phases are an ordered arc through the work: hunt for a painful problem, pick the strongest one, design the business, build the brand, build the thing itself, produce the launch material, try to break it, then package it all up. The trick is one line attached to that list: it's a floor, not a ceiling. In plain terms — this is the minimum shape the work should take, and if the agent sees a step worth adding, it should add it. You give structure without turning the brief into a cage.

The deliverables name the concrete artifacts you expect to hold at the end. The landing page. The working product. The business plan. The market research. The brand guidelines. And the connective tissue: one recap document that links to all of it. If you don't name the artifacts, you tend to get an essay about a business instead of the business.

The definition of "done" is the part that does the heavy lifting, and the part most people leave out. It's an objective bar a stranger could check — not your private sense of "good enough." Here it was: a stranger can open the recap, understand the business, watch the video, open the site, and run the demo. Read that again and notice what makes it work: someone who wasn't in the room can verify each point without asking you. That's the whole property you're after. Without a done line, an agent has no idea when to stop. It either quits while the work is half-built or polishes the same corner forever, because nothing tells it the finish line has a location.

The "never ask" rule is the line that finally lets you close the laptop: make every call yourself, write down why you made it, and don't come back until the definition of done is met. It reads like the boldest instruction in the brief, but it's the most earned. It only works because the guardrails already made every unsupervised decision safe, and the done line already drew the finish. Hand an agent "never ask me anything" without those two in place and you don't get autonomy — you get confident nonsense, delivered fast.

Orchestration: tell it how to work, then let it go further

There's a second layer to a strong brief, above the what: a few words on how the agent should work. Not step-by-step instructions — patterns. You hand it a handful and, again, call them a floor rather than a ceiling.

Fan out parallel researchers across different sources and angles, so a single blind spot in one place doesn't quietly sink the whole picture. Run an idea tournament: independent agents pitch competing directions, a panel of judges scores each one against plain criteria — how painful the problem is, how urgent, whether people would pay, how hard it is to build — and only the strongest survive. Verify adversarially: for every claim that matters, put a skeptic on it whose only job is to tear it down. A claim that lives through a real attempt to break it is one you can actually lean on. And close each phase with a completeness critic — an agent that checks nothing was left half-finished before the work is allowed to move on.

Abstract coordination diagram: one brighter central node hands tasks out to many smaller ones, a few send back a dissenting signal, and it all converges again; a restrained palette of graphite, blue and green.
Abstrakcyjny schemat koordynacji: jeden jaśniejszy centralny węzeł rozsyła zadania do wielu mniejszych, część zwraca sygnał sprzeciwu, całość zbiega się z powrotem; powściągliwa paleta grafitu, błękitu i zieleni.

Underneath all of that sits a quieter idea worth understanding, because it changes the economics. The most capable model doesn't have to do the building. It can act as a manager instead — plan the work, delegate the actual labour to cheaper, faster models, review what comes back, and send it back when it isn't good enough. The expensive model thinks and judges; the cheaper ones do the volume. Counter-intuitively, that arrangement often lands both better and cheaper than one expensive model grinding through everything alone, because the loop of delegate, review, retry catches mistakes a single straight-through pass sails right past.

What it doesn't do, and the one principle that carries

I'll be straight with you, because the hype around this deserves a correction. It isn't one-click magic, and two limits matter.

The first: a single run gets you maybe halfway to great, not the whole way. The honest expectation is a strong first draft you then iterate on — a second run, fed the specific gaps the first one left behind, gets you much closer. Treat the first pass as raw material with good bones, not a finished business you can ship as-is.

The second matters more, and it's the one I'd fix first. The highest-leverage move sits before the build, not during it: stress-test the idea itself. A capable agent will faithfully build whatever you point it at, including a weak idea, and it will build it beautifully. That's exactly the trap. Spend your own judgement on whether the thing is worth making at all before you turn the agent loose to make all of it.

Which leaves the principle I'd actually keep. The quality of an autonomous run is capped by the quality of the brief and the definition of done — not by the model, and not by how clever the opening prompt sounds. The skill worth building isn't prompt-whispering. It's specifying an outcome precisely enough that a stranger could tell whether you hit it. Get that right and "go, and don't come back until it's done" stops being reckless and becomes the strongest instruction you can hand a machine.

So next time you catch yourself about to type "build me X," stop and write the done line first: what, exactly, would a stranger check to agree it's finished? If you can't answer that yet, the agent can't either.

Test yourself

Five questions to check what stuck from the piece on handing an agent work you can walk away from.

  1. What, according to the article, most tightly caps the quality of an autonomous run?

  2. Why is it better to state the mission as an outcome than as a list of steps?

  3. What separates a good definition of done from a weak one?

  4. When does the "decide for yourself, don't ask" rule work instead of turning into confident nonsense?

  5. What is the counter-intuitive advantage of a managing agent that hands work to helpers?