SkillBambooMenu

Articles · October 8, 2026

How to Write AGENTS.md: Short Rules, Better Coding Agents

How to Write AGENTS.md: Short Rules, Better Coding Agents

How to write AGENTS.md that stays short: keep permanent rules only, load task context on demand, and steer coding agents with questions instead of orders

Arnold Explorer

More context doesn't make a coding agent smarter. Often it does the opposite, and that changes how to write AGENTS.md.

Your AGENTS.md started as ten tidy lines. Now it's a sprawling document, the agent also reads an old architecture doc and its own memory, and the output keeps getting worse: it follows a rule you dropped months ago, mixes two patterns, and when you correct it, it apologizes and breaks something else.

Here's what belongs in the permanent instruction file, how to feed task context on demand, and how to phrase corrections so the agent fixes the real problem instead of the line you pointed at.

Why more context can make your coding agent worse

The default advice is to give the agent everything: the README, the architecture docs, the logs, the whole repository. It sounds reasonable. But more context does not automatically mean better understanding.

Every extra file is another chance for noise, a contradiction or an assumption that used to be true.

Missing context and stale context fail differently. If information is missing, the agent is uncertain and may ask or hedge. If information is outdated, the agent becomes confidently wrong — it applies an old rule with full conviction and implements the wrong thing. That makes stale context the more dangerous of the two.

There's also a mechanical limit. A larger context window lets the agent access more, but it doesn't promise equal attention to every piece of it. Bury the one constraint that matters under long logs and old discussions, and the agent has a harder time surfacing it at the right moment.

Lost in the Middle: How Language Models Use Long Contexts — How language models lose information buried in long contexts

So the real question isn't whether everything fits into the prompt. It's whether the agent actually finds the right context when it needs it. In other words, the target is minimum sufficient context: enough to make the right decision, and nothing that competes with it.

Cheat sheet: what goes in AGENTS.md, how to load task context and how to steer the agent
Cheat sheet: what goes in AGENTS.md, how to load task context and how to steer the agent

How to write AGENTS.md: permanent rules only

The most useful split is between two kinds of context. Permanent context is what stays true across almost every task: style conventions, the layering of the system, security policy, how things are named. Task context is what matters only now: one bug report, one feature request, one module, one set of logs.

Keeping both in one giant blob makes the agent's reasoning harder, because it has to work out on every task which half applies.

AGENTS.md is the home for the permanent half only. And it should stay short. If your file has grown to 5,000 lines, chances are your own developers don't read it carefully either. So there's little reason to expect the agent to weigh every line correctly.

How Claude remembers your project — Claude Code Docs — Official guidance on CLAUDE.md, AGENTS.md size and conflicts

A good permanent rule is short, concrete and hard to misread. For example:

- Do not add dependencies without approval.
- Run the test suite before reporting a task as done.
- Never change authentication code unless the task asks for it.

Each line can be checked with a yes or no. Compare that with a paragraph about “preferring clean abstractions where appropriate,” which the agent will interpret differently every time.

Here's a useful test for every line: would this still be true next month, for any task? If not, it's task context or a temporary note, and it belongs somewhere else.

What belongs in AGENTS.md and what to keep out of it
What belongs in AGENTS.md and what to keep out of it

Give context a source and an expiration date

Two habits keep instruction files and memory from quietly going stale.

First, label where each piece of context comes from: current code, AGENTS.md, a decision record, an incident note, the agent's own memory. If the agent can see the source, it can weigh it. Current production code should normally beat a planning document written two years ago. But without labels, everything looks equally trustworthy.

Second, treat temporary context as something that expires. Migration rules, incident workarounds, old feature flags and one-off notes are useful for a while and harmful afterwards. Anything the agent keeps in long-term memory needs an owner who decides when it no longer applies.

Think of instruction files and memory as code: you create them, update them and retire them, and you review them along the way.

Stale rules are the quiet failure

An outdated rule doesn't produce an error message. The agent just follows it, confidently, and writes code for a system you no longer have. So put a date or an owner next to every temporary rule, because that's the only way someone will find it and remove it later.

Load task context on demand: ask the agent what it needs

Instead of loading the whole repository and every doc at the start, begin small. State the task, give the core constraints, point to the obviously relevant files. Then let the agent request more as the task proves it needs more. This is progressive disclosure applied to coding agents.

The simplest way to do it is to ask before any change is made:

Before you change anything:
- What information do you still need?
- Which files do you think are relevant?
- What are you currently assuming?

The agent then tells you what it's missing, which beats guessing and adding files blindly.

On larger codebases, it's worth adding a second step before implementation:

Review the context you have. List any conflicting instructions,
outdated assumptions or unclear sources of truth.
Don't write code yet.

The point is to make conflicts visible while they're still words, not after they've become a new pattern in your codebase. Resolve what the agent finds, then let it implement and verify.

A progressive way to load task context instead of dumping the whole repo
A progressive way to load task context instead of dumping the whole repo

Why barking fixes triggers the apology spiral

The second half of the problem starts once the agent is working. You spot a bug and type a direct order: change this line, add that check. The model responds with an eager apology and a patch.

The patch is usually local. A plain assertion is treated as a constraint to satisfy, so the model edits the immediate spot without rethinking the system around it. The fix breaks something nearby, you correct again, and the cycle repeats. A few rounds later the context is full of apologies and the code is wrapped in defensive layers nobody asked for.

The more expensive failure is when your instruction is wrong. If you order the agent to close a resource that's already handled elsewhere, it will rarely push back. It will thank you and add the change, injecting a regression into code that was working. A direct order assumes you already know the answer — and the agent is tuned to agree with you.

Towards Understanding Sycophancy in Language Models — Why AI assistants tend to agree with users

But there's a cost on your side too. To write an exact order, you have to trace the bug and design the fix in your head first. At that point the agent is just typing, and you're doing the hard part.

Why barking fixes traps the agent in a loop of apologies and patches
Why barking fixes traps the agent in a loop of apologies and patches

Nudge with questions instead of orders

A question works differently from a statement. Because it contains an open variable, the model can't simply agree with you and has to trace what actually happens before it can answer.

Think of it as a read with a conditional write. If the problem is real, the agent finds and fixes it. If the code is already correct, it explains why, you reply “ok,” and nothing breaks.

One veteran developer, Mark Hedges, reported that just switching to phrasing every prompt statement as a question noticeably improved the agent's results. There's also a learning side: orders cap the agent at what you already know, while questions leave room for it to point out something you missed.

A practical working model is mentoring a bright but anxious junior developer. You don't dictate the exact line to type, and you don't just tell them their code is broken. You ask them to walk through a scenario until they see the problem themselves.

Instead of this order Ask this question
“Unsubscribe in the cleanup method.” “If the user leaves the screen mid-request, where does this subscription end up?”
“Add a lock here.” “What happens if two calls hit this function at the same moment?”
“Check for null on line 14.” “When could this object be uninitialized at the time the callback runs?”
“Your caching is broken.” “What does this cache return right after the underlying record changes?”

So turn your gut feeling that something is off into the question behind it. It takes a few seconds, and the agent does the tracing.

Rewriting direct orders as open questions the agent has to work through
Rewriting direct orders as open questions the agent has to work through

Keeping diffs small without “DO NOT TOUCH” rules

Another familiar failure: you ask for a three-line fix and get a diff that refactors half the module. The instinct is to write “DO NOT TOUCH OTHER CODE” in capitals.

That can backfire. A negative order still points the model at “other code,” drawing that part of the project into what it's working on.

Scolding after the fact is worse. If you complain about an oversized diff, the agent rushes to roll everything back — and in that rollback it can revert the very fix you wanted, leaving broken references behind.

The question-based alternative is a short pre-mortem (a quick look at what could go wrong, before any work starts) at the end of the prompt. Mark Hedges appends a version of this to every request:

If you stay focused on this fix and set aside other issues you notice,
would a small, targeted change be less likely to cause regressions?

The agent weighs the link between the size of a change and the risk of regressions, and arrives at a narrow diff on its own.

Rules versus reprimands

Short prohibitions in AGENTS.md, such as requiring approval for new dependencies, are fine: they're standing project policy. The pattern to avoid is the angry in-task order. During a task, steer with questions.

Key terms

  • Minimum sufficient context — The smallest set of information an agent needs to make the right decision for the current task.
  • Progressive context — Starting with a small context and adding files only when the task shows they are needed.
  • Permanent vs. task context — Permanent context holds rules that are almost always true, such as coding standards; task context holds material relevant only to the current job.
  • Context provenance — Labeling where each piece of context came from so the agent can weigh current code above old documents.
  • Apology spiral — A loop in which barked corrections make the model apologize, patch locally and break other code, prompting more corrections.
  • Socratic nudge — Steering an agent with an open question about a scenario instead of telling it exactly what to change.
  • Socratic pre-mortem — A question asked before work starts that leads the agent to conclude a small, focused change is safer.

Summary: an AGENTS.md best practices checklist

What matters is the quality of context, and quantity often works against it. A short permanent AGENTS.md, task context loaded on demand, and corrections phrased as questions give a coding agent less to misread and more room to get things right.

So start by cutting your instruction file down to rules that are always true, and the next time something breaks, change how you ask for the fix.

  • AGENTS.md holds only permanent, short, checkable rules.
  • Task context is loaded per task, not stored in AGENTS.md.
  • Every temporary rule has an owner or a date, and someone actually reviews it.
  • Before changes, the agent states what it needs and lists conflicts.
  • Corrections are questions.
  • Prompts end with a narrow-scope question instead of “do not touch.”

Check yourself

1. Your AGENTS.md has grown to thousands of lines, including old migration notes. What is the best first move?

  • A. Add a line telling the agent to ignore outdated sections
  • B. Keep only short permanent rules and move task- or incident-specific notes out
  • C. Switch to a model with a bigger context window
Show answer

Answer: B. Permanent context should be short; temporary notes belong to tasks and should expire.

2. An old architecture doc contradicts the current code. What is most likely to happen if both are in context?

  • A. The agent will ask you which one is correct by default
  • B. The agent will always trust the code
  • C. The agent may follow either source or blend them into a new pattern
Show answer

Answer: C. With two sources of truth and no provenance, the agent has to guess which to trust.

3. You suspect a subscription leaks when the user navigates away. Which prompt is most likely to get a robust fix?

  • A. “This code has a memory leak.”
  • B. “What happens to this subscription if the user navigates away while a request is in flight?”
  • C. “Cancel the subscription on line 42.”
Show answer

Answer: B. An open question makes the agent trace the scenario itself instead of patching a line to comply.

4. The agent changed a dozen files while fixing a three-line bug. What is the better approach next time?

  • A. Add “DO NOT TOUCH OTHER CODE” in capital letters
  • B. Scold it afterward and ask it to revert everything except the fix
  • C. Ask up front whether staying narrowly focused would make regressions less likely
Show answer

Answer: C. A pre-mortem question lets the agent commit to a small diff before it starts, avoiding a destructive rollback.

5. You start a new bug fix. Which approach follows the minimum-sufficient-context idea? (Select all that apply.)

  • A. Give the core constraints and ask the agent what information and files it needs
  • B. Ask the agent to list conflicting instructions and outdated assumptions before changing code
  • C. Load the whole repo, all docs and memory, then ask for the fix
Show answer

Answer: A, B. A — The agent identifies what is missing, and you load only relevant files. B — Surfacing conflicts first keeps them from turning into code.

About the author
Arnold Explorer

I build small online businesses, mostly the unglamorous kind: directories and niche sites that quietly pay the bills while bigger ideas burn cash. When someone shares a success story, I take it apart to see what they actually did and which numbers don't hold up. There are always a few.