10 Comments
User's avatar
Reed Law's avatar

The question with skills or AGENTS.md is: does this information belong outside the model? The taste skill (https://github.com/Leonxlnx/taste-skill/blob/3c7017d636c3a4aad378433ea6d0cfa6c921da4a/skills/taste-skill/SKILL.md) is mostly context noise. Policy and instruction competition can degrade performance compared to the base model. I use no skills and a minimal AGENTS.md, yet agents routinely ignore instructions. Better models, smaller and higher-quality context, and external checks are more reliable than trying to add post-training in Markdown. What belongs in AGENTS.md or skills then? Only what the model couldn't reliably know through training: specific, local policies or procedures about how you run your project.

Michael Mangialardi's avatar

I ,too, am following a minimal approach for these reasons. It does seem that higher-quality context has more value than skills and personalities. I am still in the process of evaluating this. I’d be curious to hear about your optimizations in this area.

One potential solution is to get the value of skills and personalities without the risks is to use “just in time” (JIT) optimizations. You could leverage a low-cost model to arrange the optimal skills and personalities based on the current context and prompt. This would give you the benefits without the maintenance. Perhaps this is a poor assumption. LLMs may not be as self-aware on their own potential to capture it well in skills and personalities, drifting to popular solutions rather than the ideal. I suppose this could be mitigated by managing its awareness via emergent papers, etc.—but that reintroduces the maintenance risk with unclear gains.

Reed Law's avatar

My AGENTS.md primarily describes workflow (local clone of production data; investigate/fix locally; deploy after verifying; follow post-deploy checklist). Then there's writing guidance for documentation which grew organically, mostly to combat overly verbose output. I find code conventions, such as "Always prefer existing patterns over creating new ones", do little. "Skills" is a misnomer because it suggests added capability when in fact they add context. "Personalities" are also at the context layer. They do steer tendencies like conciseness. The most effective JIT-like tool is retrieval. I have workflows that do something like what you describe, except the middle layer is assembled context from database results (RAG). For example, a scanned document is parsed by one model that outputs named entities which are used to query the database. The query results are formed into a prompt which is used by another model. That way, both the document and database context is in view. Coding harnesses do essentially the same thing: parse the user's request, query the codebase using grep or AST, assemble results into context, and then ask the model to work with that context. A detailed style guide in a project may be more useful than generic skills because it is locally maintained, and AGENTS.md can route design decisions to it. Skills, on the other hand, first have to be discovered. Codex has an open issue (https://github.com/openai/codex/issues/32101) for degraded discovery of deferred MCP tools. That illustrates the problem with skills. When context is hidden, you're dependent on the harness getting the routing right.

Michael Mangialardi's avatar

Interesting, and you're right that skills are effectively context. The only path to optimization seems to be to have benchmarks that A/B test the impact of context. When I have run tests, sometimes skills to improve prompts worsen the result.

There is another axis to consider: generative vs. deterministic.

Knowing when you need generation (e.g., creating a shared library) versus glueing to something determined (e.g., connecting to a shared library) is one axis, and the other is context. Ultimately, you are trying to get the best result in the most efficient manner, and pointing an LLM to consume an external library may be more straightforward than optimized context.

I suppose a third axis is "worth reasoning about" and "eh, it will get there eventually."

Reed Law's avatar

I'm not sure whether context should be prioritized over model selection when the latter is much higher leverage. I've observed remarkable failures in basic instruction following using the latest models. Given two models with the same AGENTS.md, if one follows it more consistently than the other, that alone may dominate any other gains to be had through context optimization.

Michael Mangialardi's avatar

Ultimately, it all comes back to the need for benchmarks on how the model is working before introducing abstractions.

As you note, do your AGENTS.md lack context, or did the model make a mistake?

Is enriched context moving the needle with respect to time, token consumption, or output? Benchmarks per model, per domain are the answer.

Reed Law's avatar

I ran an experiment by asking four models across three agents to pick up the same to-do item in a design document. Then I stashed the resultant worktree from each and asked another model to grade 1-10 on instruction following, task completion, code quality, and big picture understanding. I allowed models to judge themselves, but also included one judge (gemini-3.1-pro-preview) that was not a participant in order to mitigate bias. All models selected claude-opus-5-high as the winner, even codex-gpt-6-astra-low. What's more interesting is that Astra identified the same issue with Opus that has been plaguing me (overly verbose reporting of investigation findings). The Opus issue persists despite multiple AGENTS.md counter-prompts, which suggests the model itself has a tendency to zealously report details even when already addressed. I normally don't take the time to do tests like this, but the Claude issue was grating enough that I became open to switching. I have been regularly switching between Claude and Codex because they both have their strengths along with frequent regressions. The two open weight models I tried (GLM and Minimax) both failed miserably, although I used another harness (OpenCode) with them, so it's hard to judge fairly.

Craig Smitham's avatar

It depends how you're using skils/context. There is a way of making skills/instructions that just micro-manage the model. Sometimes that can help less-intelligent models. That's not the right way to use them. Where instructions/context have their value is in the consistent grammar/world/system they represent. They are like knives though. A dull knife is dangrous/harmful. A sharp knife needs to be maintained and tested.

Craig Smitham's avatar

Another way to look at it: if you see skills/instructions as a way to extend the power/language of the prompt/prompter, not as a compensation mechanism for the model.

Michael Mangialardi's avatar

Good thoughts. The trick is knowing when language transformation/extension dulls the knife or sharpens it, and then knowing the reusability of that sharpening-effect when applied to different domains. One solution is a JIT approach where the skill doesn’t autorun each time but a quick check (via an LLM) determines the ideal skills. I believe Cursor is moving in this direction.

When tools like Cursor have a window into the evaluation, it can help strengthen the case again a DIY approach.