When Agent Docs Lie: Fixing Skill and Config Drift in Public

Agents do not improvise your process. They read the skill folders, blueprints, and agent markdown you put in front of them - then they act. When those files drift from the code that already shipped, the agent does not get confused. It gets confidently wrong.

In September 2026 a public pull request made that failure mode impossible to ignore: agent-facing docs in an open repo still told agents to reject a rule that production had already accepted. The fix was not a clever prompt. It was correcting the SoT and putting the correction where every harness can install it.

This is the pattern - and the system we recommend when teams run Claude Code, Codex, Cursor, or all three.

PR #109 Public drift fix

rmems/grok-ozempic merged a docs correction after agent config inverted a shipped V2 rule

~17k Agent sessions studied

Armature compared Claude Code, Codex, and Cursor - they agree on the same tool in only ~42% of cells

1 SoT Many installs

Skill packs are meant to be versioned folders - runtime dirs should be copies, not separate edit homes

The Situation

Shared skills are now shipping the way software ships: git repos and skill folders. Kits like NeoLabHQ's context-engineering pack are built to install across Claude Code and related harnesses. That reuse is the point - and it is also how a stale sentence becomes a fleet-wide bug.

The grok-ozempic case was blunt. Agent config still said the runtime V2 structural manifest "is rejected until GH #40" - after #40 had already shipped. A session reading that blueprint was steered toward the old baseline and away from the fail-closed path built specifically to stop silent ternary quantization of routers and norms.

What was actually broken

  • Docs contradicted production. The agent-facing rule was the opposite of the shipped V2 acceptance path.
  • Trust amplified the damage. Agents treat repo skills and blueprints as law, not as suggestions.
  • Secondary drift compounded. Related files still pointed at closed issues, incomplete test lists, and dead commands - each cheap alone, expensive together.
  • Install surfaces multiplied risk. Every harness that copied the pack inherited the lie until the SoT was fixed.
The Hidden Cost

A wrong skill file does not fail loudly. It produces a clean, confident agent run that packs the wrong defaults, skips the fail-closed path, or "fixes" problems that were already closed. You pay in rework, not in an obvious crash.

The Approach: Fix the SoT, Then Reinstall

The instinct when an agent misbehaves is to add more prompt text in chat. That patches one session. The durable fix is the same discipline as code: correct the versioned source, then refresh every install from that source.

PR #109 did exactly that - verified against the tree, not assumed:

  1. Correct the agent-facing rule so V2 is accepted and fail-closed, not "rejected until a closed issue."
  2. Align companion docs (CLAUDE/REVIEW-style files) with the real test suite and open work.
  3. Remove dead instructions that burn agent turns on commands and tickets that no longer exist.

That is the entire intervention. It takes a PR review, not a new framework.

Before vs After: What the Agent Sees

The contrast is easiest to see as what an agent is told to prefer.

Before - drifted instructions

// Agent reading stale blueprint
A
Agent (from docs)
V2 structural manifest is rejected until GH #40. Prefer the V1 baseline. Steer away from the fail-closed path.
pre-fix

After - SoT matches shipped code

// Agent reading corrected SoT (PR #109)
A
Agent (from docs)
V2 is accepted and fail-closed. Prefer structural-manifest.json. Do not chase closed issues or dead commands.
post-fix
What Actually Changed

The model did not get smarter overnight. The instructions stopped lying. Same as training a coordinator on Claude: the leverage is a shared, visible system - not heroic one-off chats.

Why Multi-Harness Teams Feel This First

Armature's public study of nearly 17,000 coding-agent sessions shows Claude Code, Codex, and Cursor already disagree on discovery sources and tool picks. That diversity is fine for choosing a database vendor. It is dangerous for your skill SoT.

Drift pattern Stable pattern
Each harness keeps a locally edited skill tree One versioned skill pack in git; installs are copies
Chat patches "fix" one agent's memory PR fixes the SoT; every runtime reinstalls
Docs describe aspirational future ("until issue X") Docs verified against the tree that shipped
Closed tickets and dead CLIs still in agent guides Agent guides only list live commands and open work
Assume all harnesses read the same web priors Assume installs diverge - SoT must be explicit

The model is not the source of truth. Your skill pack is. If the pack lies, every harness that installs it will lie politely.

What Good Looks Like

Public correction, not quiet folklore Drift fixed in a reviewable PR with before/after tied to real files - see grok-ozempic #109.
Skills treated as products Packs like context-engineering-kit are installed artifacts - version them, don't hand-edit one runtime's copy.
Harness differences acknowledged Claude Code, Codex, and Cursor will not always pick the same tools; your SoT must not depend on them sharing one brain.
Fail-closed agent rules When production ships a stricter path, docs must promote it the same day - not preserve the old escape hatch.

The Repeatable System

You do not need a new platform. You need a boring loop that matches how you already ship code.

When you change agent behavior

  • Edit the versioned skill / blueprint / agent markdown in git
  • Verify each claim against the tree (tests, flags, open issues)
  • Open a PR - especially when the old text inverted a shipped rule

When you install into a harness

  • Copy or install from the SoT pack - do not fork by accident in a local skills folder
  • After SoT merges, refresh Claude Code / Codex / Cursor installs the same way you pull main
  • Delete or ignore stale local copies that were hand-patched in chat eras

When an agent "goes weird"

  • Ask which file it is obeying before you blame the model
  • Diff that file against production reality
  • Fix SoT first; only then re-run the agent
Who This Works For

Any team running more than one coding agent, or shipping shared skills to clients and staff. If two people can install different skill trees, you already have a drift surface.

Sources

Build Agents That Obey the Right Rules

We help teams design agent systems and training so skills, configs, and workflows stay aligned with what actually shipped.

See AI Agent Systems →
AB
AiBrainBuilders Team
AI Agent Builders & Trainers

We build AI agents for businesses and train the teams that run them. Every post comes from real build experience - things that worked, things that didn't, and the decisions that made the difference.