Plan drift: when a merged PR quietly undoes a team decision
Decision drift is when merged code quietly contradicts a team decision. See how it happens and run a three-step drift check this week, no tooling needed.
On this page
Decision drift is what happens when merged code quietly contradicts a decision your team made somewhere else, usually in a Slack thread. CI is green, the review was reasonable, and nobody notices that the plan changed. You can catch it without new tooling: write each decision where agents and reviewers read it, link the implementing PRs back to it, and run a short weekly pass that compares decisions to merges.
What is decision drift, and how does it show up?
Your agent-written PRs pass CI, but nobody knows if they match the plan.
An illustration with sample data: on Tuesday, a thread settles that exports will stay synchronous for now, because the async queue is out of scope this month. Four people react with a thumbs up. On Friday, an agent opens a PR titled "Improve export reliability" that adds a background job and a retry queue. The description is accurate. The tests pass. The reviewer, who was not in Tuesday's thread, sees a sensible change and approves it. It merges.
Nothing looks wrong. The following Wednesday someone asks why there is a queue, and half the team says they thought that was off the table. Now you are arguing about what was agreed, with no record that settles it.
That is decision drift. The code is not broken. It just no longer matches what the team decided, and nobody chose that on purpose.
Why does this keep happening, and what does it cost?
AI-native teams create context faster than people can sync, manage or remember it. Agents open PRs all day, decisions get made in threads, and plans shift mid-week. The rituals that used to keep everyone aligned (standup, status report, sprint planning) were built for human-speed change. When the code moves at agent speed, the decision and the diff stop meeting. It is the same gap that makes it hard to keep an AI-native team on the same page.
At Liouville Labs we run about a dozen products with agents writing most of the code, so we see this pressure on our own team. The mechanics are simple:
- The decision lives in Slack. The agent never saw it, and the reviewer may not have either.
- Each PR is locally reasonable. Read alone, it makes sense.
- Volume hides the contradiction. A reviewer with a long queue checks whether the change works, not whether it was supposed to exist.
The costs are plain. You rework code after the merge went the wrong way. A decision gets silently reversed, so the next person builds on the wrong assumption. Or the team reopens a settled question because nobody can show it was settled. All three eat time you already spent once.
What do people call this?
Several names circle the same problem, from slightly different angles:
- Plan drift: the merged work no longer matches the plan the team agreed on. This is the product leader's view, and the term we use here.
- Decision drift: the specific decision, often made in chat, is contradicted by later changes. AsDecided writes about it in the context of AI-assisted development.
- Spec drift: the implementation wanders from the written spec. SpecStory covers it under that name.
- Architectural drift: the codebase slowly departs from its intended structure. Mneme uses this term.
The difference is mostly where the source of truth sits: a spec, an architecture, or a conversation. Fresh product calls usually live in conversations, which is why they drift first.
What do teams usually try, and where does it fall short?
CLAUDE.md files and ADRs. These are good for durable technical rules: the framework we use, the naming scheme, the pattern for error handling. Agents and humans both read them. They miss fresh product calls. Nobody writes an ADR for "exports stay synchronous this month," and by the time someone would, the PR has merged. (If you are weighing the two formats, see ADRs or a decision log.)
A PR template that asks "link the decision." Good intent, and cheap. Under volume it becomes a field people fill with "n/a" or skip. An agent filling the template has nothing to link to if the decision was never written down.
Careful review. The right instinct, and nothing replaces a human on a risky change. It does not scale when agents open more PRs than anyone can read, and a reviewer cannot check a diff against a decision they have never seen. We cover that volume problem in too many PRs to review.
A decision log. The strongest of the usual options, and the base of the fix below. On its own it fails when the log and the PRs never touch each other. See the decision log template for Slack for how to write the entries.
What works instead: a three-step drift check you can run this week
The goal is not to review every diff. It is to make decisions visible to the people and agents writing code, and to compare decisions to merges on a schedule. Three steps.
Step 1: Write the decision where agents and reviewers read it
When a thread settles something, one person writes a short entry the same day. Keep it to four lines:
- What was decided, in one sentence, with the verb in it ("Exports stay synchronous through October").
- What it rules out ("No background queue for exports").
- A link to the thread.
- The date, and who can change it.
Put it in a file in the repo (next to CLAUDE.md, or a DECISIONS file) if it constrains code. Agents that read repo instructions will see it, and reviewers can find it from the diff. If it is a product call with no code rule, keep it in your decision log but still link it from the relevant issue or project doc.
The "what it rules out" line matters most. Drift usually happens in the gap between what was chosen and what was excluded.
Step 2: Link implementing PRs back to the thread
Two directions, both cheap:
- In the PR description, name the decision it implements or touches. For agent PRs, add that to the task prompt: "This work is governed by the decision in this file. If the change conflicts with it, stop and say so."
- In the decision entry, add the PR links as they merge.
This gives you a trail. When someone asks "does this match what we agreed?", the answer is two clicks away instead of an afternoon of searching. It also makes a missing link informative: a PR in the area of a decision with no link is worth a second look.
Step 3: Run a weekly 10-minute decisions-vs-merges pass
Pick a fixed slot, ideally Friday or Monday. One person (rotate it) takes the decisions made or touched this week and scans the merged PRs in the same areas. You are not reading code line by line. You are reading titles, descriptions and changed-file lists, and asking one question per decision: did anything merge that contradicts this?
Hypothetically, say a week has 60 merged PRs and five new decisions. You do not read 60 PRs. You filter to the files and areas the five decisions cover, which is usually a handful, and check those.
Use this checklist each week:
- List every decision made or changed this week, with its "rules out" line.
- For each, filter merged PRs by the area or files it covers.
- Check each hit: does the title, description or file list conflict with the decision?
- For PRs with no decision link in a decision area, open the thread and confirm they are consistent.
- For any conflict, write down one of two outcomes: revert or fix the code, or record a new decision that supersedes the old one.
- Update the decision entry with the PR links and the outcome.
- Note any decision made this week that never got written down, and write it now.
The last two outcomes in the list are the point. Drift is only a problem when it is silent. If the team looks at a contradicting merge and says "actually, the queue is right," that is a legitimate new decision. Write it down so the record stays true.
Where does Biddle fit?
Biddle keeps decisions with evidence (the PR, commit or thread behind each one) and is learning to flag a merged PR that contradicts one. That flagging is not shipped, so treat the weekly pass above as the real mechanism for now. The principle is to stay quiet unless something matters, because a wrong nudge costs more trust than silence (see nudges and the ground rules). Biddle is in early access, reads GitHub and Slack only, is read-only on your repos, and runs every weekday for our own team, and you can see what its output looks like in a sample briefing. If decision drift is a problem on your team, you can request early access.
Common questions
Is decision drift the same as a bug?
No. A bug is code that does not do what it was meant to do. Decision drift is code that does exactly what its author meant, but contradicts what the team agreed. Tests pass either way, which is why it slips through.
Does this only happen with agent-written code?
No, people drift from decisions too. Agents make it more likely, because they produce more PRs, work without having sat in the thread, and write convincing descriptions that make a contradicting change look routine.
How many decisions should we track?
Only the ones that rule something out and that someone could plausibly contradict by accident. A short list you actually check beats a complete one nobody opens.
Read next
- The decision was made in a thread nobody can find: a decision log template for Slack: how to write the entries Step 1 relies on.
- ADRs or a decision log: which one a small product team needs: where durable technical rules end and fresh product calls begin.
- Too many PRs to review? What to track when agents ship 200+ in two weeks: what to watch when reading every diff stops working.
Want this in your Slack?
Biddle reads your GitHub and Slack and sends you a briefing each weekday morning, with every claim linked to the PR or thread behind it. We’re onboarding a few product teams now, and we set each one up with you.
Request early accessRelated posts
ADRs or a decision log: which one a small product team needs
ADRs hold technical decisions agents must follow; a decision log holds product calls made in Slack. See them side by side, and make both agent-readable.
AI answers in Slack: who sees what, and why private channels stay private
How an AI assistant in Slack should scope answers to private channels, the visibility rule Biddle uses, and a checklist of questions to ask any tool first.
Decisions lost in Slack threads: a decision log template
A decision log template for Slack: one row per decision with thread permalink, implementing PRs and superseded status, plus a one-minute capture habit.