Skip to content

Too many PRs to review? What to track when agents ship 200+ in two weeks

When agents open more PRs than anyone can read, stop tracking diffs. Track decisions, open questions and needs-a-human in a 15-minute routine.

8 min read

  • agent-era teams
  • trust
On this page

If you have too many pull requests to review, the fix is not reading faster. Stop treating the diff as the unit of tracking and keep three lists instead: decisions in force, open questions, and things that need a human. Then give a small, rule-based set of PRs a real human look and let the rest be covered by CI, scoped review and those lists.

The rest of this post explains why the volume problem is structural, what teams usually try, and a routine you can run starting tomorrow.

What does it feel like when agents open more PRs than you can read?

Say 40 PRs merged while you were out. You open the GitHub tab, sort by recently merged, and scroll. Titles are plausible. Descriptions are tidy. CI is green on nearly all of them. You can read perhaps five properly before your next meeting.

So you do what everyone does: skim titles, open two or three that look risky, and close the tab. Then, on Thursday, someone asks whether the new pricing flow still follows what you all agreed on Monday, and you realize you cannot say. You are still expected to know what changed. You just have no way to know it.

This is the moment to notice: the problem is not that you are slow. The reading was never going to fit.

Why does this keep happening, and what does it cost?

AI-native teams create context much faster than people can sync, manage or remember it. Agents open PRs continuously, decisions get made in Slack threads, and plans change mid-week. Rituals built for human-speed change, like the standup, the weekly status report and the review queue, assume the amount of change fits in a person's head. When the code moves at agent speed, it does not. (We cover the wider version of this in keeping an AI-native product team on the same page.)

We see this at Liouville Labs. We run about a dozen products with agents writing most of the code, and one product shipped 200+ PRs in two weeks. In the 14 days to October 3, 2026, that product merged 197 PRs, and its busiest 14-day window since August 1 reached 295. Nobody reads that many diffs and also does their own job.

Third-party data points the same direction. LinearB's 2026 benchmarks report that AI-generated PRs wait about 5.25x longer to be picked up for review than unassisted PRs (see LinearB's 2026 benchmarks on AI PR merge rate). There is also an arXiv study of agentic pull requests if you want the research view of this shift. Output scales with compute. Review attention scales with people.

The cost shows up in three places:

  • Time reassembling context. You rebuild the picture from PR titles, threads and memory every time you return from a day away.
  • Decisions missed or reopened. A choice made in a thread never reaches the PR that touches it, so it gets made again, differently.
  • Rework found late. A merge that quietly went the other way is discovered after other work has built on top of it. We cover that pattern in plan drift: when a merged PR quietly undoes a team decision.

What do teams usually try, and where does it fall short?

Reading every diff. This is the right instinct for high-risk code, and nothing replaces a careful human read where it matters. It breaks on volume: at hundreds of PRs a fortnight, the careful read becomes a skim, and a skim gives false comfort.

Per-PR Slack notifications. Good for getting a reviewer's attention on one specific PR. At agent volume the channel becomes a feed nobody reads, and the important item looks identical to the routine one. We compare this approach with digests and briefings in your GitHub channel posts every event and nobody reads it.

Merge digests. A daily list of what merged is better than a stream, and it is cheap to set up. But a list of titles is still a list of diffs. It tells you what changed, not what it means for the plan.

Metrics dashboards. Cycle time, review wait and merge rate are useful for spotting a queue problem. They cannot tell you whether a merged change matches what the team decided. A fast, green merge can still be the wrong merge.

What these share: they all organize information around the PR. The thing you actually need to track is not the PR.

What works instead: track three lists, not diffs

The unit that matters to a product leader is the decision. A PR is just evidence that a decision was or was not followed. So track three lists and let PRs hang off them as links.

An illustration with sample data, for a team shipping a checkout change:

ListWhat goes inSample entry (illustrative)Evidence to link
Decisions in forceChoices the team made that code should now follow"Guest checkout stays; no forced account creation"The thread where it was decided, plus PRs that touch checkout
Open questionsThings asked and not yet answered"Do we show tax before or after address entry?"The thread and the PR waiting on the answer
Needs a humanMerges or PRs where a person must look"Auth change touching session expiry"The PR, with the reason it was flagged

Keep each entry to one line plus a link. If you cannot link the evidence, the entry is a rumor, so mark it as one.

The 15-minute daily routine

Run this once each weekday morning, before standup or instead of it. An illustration of the order and a rough time budget, not a measured result:

  1. Scan merged PRs since yesterday (5 minutes). Do not read diffs. Read titles and descriptions and ask one question per PR: does this touch a decision on the list or raise a new one?
  2. Update decisions in force (4 minutes). Add anything decided in a thread since yesterday. Strike or mark superseded anything a later decision replaced. Some decisions do get replaced: in Biddle's record on our own team, 13 of the 448 decisions decided between September 1 and October 3, 2026 were superseded by a later decision. That is one team over about a month, not a benchmark, but it is why the list needs a way to mark a decision as replaced.
  3. Update open questions (3 minutes). Add new ones. Close any that a decision now answers.
  4. Pick the human looks (3 minutes). Apply the rule below, assign each to a named person, and write down why.

A rule for which PRs get a human look

Pick criteria in advance so you are not choosing by mood. A reasonable starting set:

  • The PR touches something on the decisions-in-force list, and the description does not say how it follows it.
  • It changes auth, payments, permissions, data deletion or anything hard to roll back.
  • The description and the changed-file list do not match (a small description over a wide change).
  • It resolves or contradicts an open question without citing the answer.
  • CI is green but the change is the first of its kind in that area, so there is no test history to lean on.

Everything else can rely on CI, normal scoped review and the lists. Revisit the rule every two weeks: if a miss slipped through, add the criterion that would have caught it.

Two habits that make the lists hold

  • Write decisions where the lists can find them. A decision that lives only in a thread decays quickly. Our decision log template for Slack gives you a format for one-line entries.
  • Make agents state intent. Ask that every agent-written PR description say which decision or question it addresses. It is a signal, not proof, but it gives step 1 of the routine something concrete to check against.

Where does Biddle fit?

Biddle is an AI chief of staff for product teams. It reads GitHub and Slack (those are the only sources today; Linear and Notion are coming) and sends each leader a weekday morning briefing with what needs you, what changed since the last briefing, and what is worth a look, with links to the PRs and threads behind each line. You can see the shape in the sample briefing, the reasoning in why now, and the mechanics in how it works. It is in early access and runs every weekday for our own team; it is read-only on GitHub and does not comment on PRs or write code, and flagging a merged PR that contradicts a decision is something it is still learning, not something it does today. If you want that list waiting for you instead of building it by hand, you can request early access to Biddle.

Common questions

Should I stop reviewing agent PRs altogether?

No. Keep humans on the PRs your rule selects, such as irreversible changes and ones that touch a decision in force. The goal is to spend human attention where it changes the outcome, not to remove it.

Is green CI enough to merge an agent PR?

Green CI tells you the checks passed, not that the change matches what the team decided. That gap is why the decisions list matters: it gives you something to compare the PR against.

How many decisions can one person realistically track?

Fewer than a fast team generates. On our own team, Biddle's record holds 448 decisions decided between September 1 and October 3, 2026, and 124 open questions as of October 3. That is one team over about a month, not a benchmark, but it shows why a written list beats memory.

Want this in your Slack?

Biddle reads your GitHub and Slack and sends you a briefing each weekday morning, with every claim linked to the PR or thread behind it. We’re onboarding a few product teams now, and we set each one up with you.

Request early access

9 min read

What is an AI chief of staff for product teams?

What an AI chief of staff for product teams does, how it differs from email assistants and exec dashboards, and how to tell whether your team needs one.

  • ai chief of staff
  • management