When Agent Docs Lie: Fixing Skill and Config Drift in Public
Agents do not improvise your process. They read the skill folders, blueprints, and agent markdown you put in front of them - then they act. When those files drift from the code that already shipped, the agent does not get confused. It gets confidently wrong.
In September 2026 a public pull request made that failure mode impossible to ignore: agent-facing docs in an open repo still told agents to reject a rule that production had already accepted. The fix was not a clever prompt. It was correcting the SoT and putting the correction where every harness can install it.
This is the pattern - and the system we recommend when teams run Claude Code, Codex, Cursor, or all three.
rmems/grok-ozempic merged a docs correction after agent config inverted a shipped V2 rule
Armature compared Claude Code, Codex, and Cursor - they agree on the same tool in only ~42% of cells
Skill packs are meant to be versioned folders - runtime dirs should be copies, not separate edit homes
The Situation
Shared skills are now shipping the way software ships: git repos and skill folders. Kits like NeoLabHQ's context-engineering pack are built to install across Claude Code and related harnesses. That reuse is the point - and it is also how a stale sentence becomes a fleet-wide bug.
The grok-ozempic case was blunt. Agent config still said the runtime V2 structural manifest "is rejected until GH #40" - after #40 had already shipped. A session reading that blueprint was steered toward the old baseline and away from the fail-closed path built specifically to stop silent ternary quantization of routers and norms.
What was actually broken
- Docs contradicted production. The agent-facing rule was the opposite of the shipped V2 acceptance path.
- Trust amplified the damage. Agents treat repo skills and blueprints as law, not as suggestions.
- Secondary drift compounded. Related files still pointed at closed issues, incomplete test lists, and dead commands - each cheap alone, expensive together.
- Install surfaces multiplied risk. Every harness that copied the pack inherited the lie until the SoT was fixed.
A wrong skill file does not fail loudly. It produces a clean, confident agent run that packs the wrong defaults, skips the fail-closed path, or "fixes" problems that were already closed. You pay in rework, not in an obvious crash.
The Approach: Fix the SoT, Then Reinstall
The instinct when an agent misbehaves is to add more prompt text in chat. That patches one session. The durable fix is the same discipline as code: correct the versioned source, then refresh every install from that source.
PR #109 did exactly that - verified against the tree, not assumed:
- Correct the agent-facing rule so V2 is accepted and fail-closed, not "rejected until a closed issue."
- Align companion docs (CLAUDE/REVIEW-style files) with the real test suite and open work.
- Remove dead instructions that burn agent turns on commands and tickets that no longer exist.
That is the entire intervention. It takes a PR review, not a new framework.
Before vs After: What the Agent Sees
The contrast is easiest to see as what an agent is told to prefer.
Before - drifted instructions
After - SoT matches shipped code
The model did not get smarter overnight. The instructions stopped lying. Same as training a coordinator on Claude: the leverage is a shared, visible system - not heroic one-off chats.
Why Multi-Harness Teams Feel This First
Armature's public study of nearly 17,000 coding-agent sessions shows Claude Code, Codex, and Cursor already disagree on discovery sources and tool picks. That diversity is fine for choosing a database vendor. It is dangerous for your skill SoT.
| Drift pattern | Stable pattern |
|---|---|
| Each harness keeps a locally edited skill tree | One versioned skill pack in git; installs are copies |
| Chat patches "fix" one agent's memory | PR fixes the SoT; every runtime reinstalls |
| Docs describe aspirational future ("until issue X") | Docs verified against the tree that shipped |
| Closed tickets and dead CLIs still in agent guides | Agent guides only list live commands and open work |
| Assume all harnesses read the same web priors | Assume installs diverge - SoT must be explicit |
The model is not the source of truth. Your skill pack is. If the pack lies, every harness that installs it will lie politely.
What Good Looks Like
The Repeatable System
You do not need a new platform. You need a boring loop that matches how you already ship code.
When you change agent behavior
- Edit the versioned skill / blueprint / agent markdown in git
- Verify each claim against the tree (tests, flags, open issues)
- Open a PR - especially when the old text inverted a shipped rule
When you install into a harness
- Copy or install from the SoT pack - do not fork by accident in a local skills folder
- After SoT merges, refresh Claude Code / Codex / Cursor installs the same way you pull main
- Delete or ignore stale local copies that were hand-patched in chat eras
When an agent "goes weird"
- Ask which file it is obeying before you blame the model
- Diff that file against production reality
- Fix SoT first; only then re-run the agent
Any team running more than one coding agent, or shipping shared skills to clients and staff. If two people can install different skill trees, you already have a drift surface.
Sources
- rmems/grok-ozempic PR #109 - correct agent-facing drift that inverted the shipped V2 rule
- NeoLabHQ/context-engineering-kit - shared skill pack
- Armature - Which tools do Claude Code, Codex and Cursor choose? (~16,893 sessions)
Build Agents That Obey the Right Rules
We help teams design agent systems and training so skills, configs, and workflows stay aligned with what actually shipped.
See AI Agent Systems →